Introducing SWE-bench Verified
We’re releasing a human-validated subset of SWE-bench that more reliably evaluates AI models’ ability to solve real-world software issues.
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
We’re releasing a human-validated subset of SWE-bench that more reliably evaluates AI models’ ability to solve real-world software issues.
arXiv:2607. 29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important.
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.
arXiv:2608. 19653v1 Announce Type: cross Abstract: Autonomous agents for machine learning experimentation must navigate heterogeneous repositories, repair training pipelines, and evaluate candidate improvements under realistic compute constraints.
Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.
The article "Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks" surveys the lack of a standard definition for AI agents and organizes this ambiguity into five dimensions: environmental interaction, learning and adaptation, autonomy, goal‑directed behavior, and temporal coherence. It reviews how each dimension has been conceptualized in prior work and compiles the metrics, benchmarks, and evaluation frameworks used to assess them. The authors also introduce the Agent Compendium, a public digital resource that extends these evaluation methods, aiming to provide a common structure for evaluating and comparing agent capabilities across AI systems.
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.
arXiv:2606. 04971v1 Announce Type: new Abstract: Machine learning engineering (MLE) agents promise to automate end-to-end ML pipeline development from raw data and natural language instructions, potentially making ML accessible to non-technical domain experts.
arXiv:2609.40303v1 Announce Type: new Abstract: Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnati...
Despite the rapid progress of LLM-based agents for planning, code generation, and debugging, their practical value for tabular machine learning remains underexplored. In this paper, we investigate a c...