arXiv AI By Pavel Adamenko, Mikhail Ivanov, Aidar Valeev, Rodion Levichev, Pavel Zadorozhny, Ivan Lopatin, Dmitry Babaev, Alena Fenogenova, Valentin Malykh

SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

Read the original on arXiv AI →

arXiv:2507. 11059v3 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench dataset.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

E2E-SWE is a benchmark that tests large language models’ ability to create complete, functional software repositories from scratch. It includes 186 tasks across 11 programming languages, each requiring an agent to build an installable project based solely on a natural‑language specification and an empty workspace, while passing a hidden test suite. The benchmark was crafted by software engineers and LLMs, then refined through iterative verification by autonomous agents to ensure clarity and solvability.

By Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou
arXiv Machine Learning
Jun 9

RepoLaunch: Automating Build and Management of Code Repositories across Languages and Platforms

arXiv:2603. 05026v2 Announce Type: replace-cross Abstract: Language model (LM) agents have driven substantial progress in automated software engineering (SWE), yet building and testing software repositories at scale remains a largely manual and labor-intensive bottleneck.

By Kenan Li, Rongzhi Li, Linghao Zhang, Qirui Jin, Liao Zhu, Xiaosong Huang, Geng Zhang, Yikai Zhang, Shilin He, Chengxing Xie, Xin Zhang, Zijian Jin, Bowen Li, Chaoyun Zhang, Yu Kang, Yufan Huang, Elsie Nallipogu, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang
arXiv AI
Sep 24

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.

By Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati
arXiv AI
Jul 22

Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality

arXiv:2607. 03691v2 Announce Type: replace-cross Abstract: Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agent harness: a middleware layer in between a developer and a large language model that orchestrates system prompts, tool execution, context management, and iterative reasoning loops.

By Oussama Ben Sghaier, Hao Li, Bram Adams, Ahmed E. Hassan