arXiv:2511. 06090v3 Announce Type: replace-cross Abstract: Optimizing the performance of large-scale software repositories demands expertise in code reasoning and software engineering (SWE) to reduce runtime while preserving program correctness.
By Jeffrey Jian Ma, Milad Hashemi, Amir Yazdanbakhsh, Kevin Swersky, Ofir Press, Enhui Li, Vijay Janapa Reddi, Parthasarathy Ranganathan
arXiv:2606. 20512v1 Announce Type: cross Abstract: LLM-based coding agents need higher-level operational knowledge about a repository (which files house which subsystems, how to run the test suite, which workflows have historically led to wrong fixes) that does not exist in the code itself.
By Asa Shepard, Jeannie Albrecht
arXiv:2507. 22580v2 Announce Type: replace-cross Abstract: Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention.
By Marcos Fuster-Pena, David de-Fitero-Dominguez, Antonio Garcia-Cabot, Eva Garcia-Lopez
arXiv:2608. 09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.
By Xin Zhou, Chun Yong Chong, Kisub Kim, Yun Peng, Rui Shu, Zihan Wu, Xu Han, Guowen Yuan, Zeyang Zhuang, Jounghoon Kim, Jeongjin Ju, Seongmin Ju, Taein Yoon, David Lo
arXiv:2607. 28587v2 Announce Type: replace-cross Abstract: SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability.
By Manyi Wang, Junjielong Xu, Pinjia He
arXiv:2607. 02370v1 Announce Type: cross Abstract: Compiler missed optimizations refer to cases in which compilers failed to optimize certain code.
By Batu Guan, Zirui Wang, Shaohua Li
arXiv:2509. 24148v3 Announce Type: replace-cross Abstract: Test-Driven Development (TDD) is a widely adopted practice that requires developers to create and execute tests alongside implementation.
By Yiran Hu, Nan Jiang, Shanchao Liang, Yi Wu, Lin Tan
The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.
By Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati
The paper investigates how large language models (LLMs) handle bug fixing versus problem solving in competitive programming. Using a dataset of ~3,000 Codeforces submissions and their human fixes, the authors compare LLM-generated patches to human patches and assess whether LLMs prefer to modify buggy code or generate new solutions. Results show that LLMs often alter more lines than necessary and sometimes produce entirely new solutions, performing better when allowed to generate solutions from scratch rather than patching existing code.
By Alexandru Stefan Stoica, Traian Rebedea, Marian Cristian Mihaescu
The paper investigates how Large Language Models (LLMs) handle bug fixing compared to human-written patches by analyzing about 3,000 Codeforces submissions. It finds that LLMs often modify more lines than necessary and sometimes produce entirely new solutions, and that they solve more problems correctly when generating solutions from scratch rather than patching existing code. The study highlights implications for AI‑assisted programming tools, suggesting a shift toward incremental problem‑solving strategies.
arXiv:2606. 12344v1 Announce Type: new Abstract: General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring.
By Mengyu Zheng, Kai Han, Boxun Li, Haiyang Xu, Yuchuan Tian, Wei He, Hang Zhou, Jianyuan Guo, Hailin Hu, Lin Ma, Chao Xu, Guohao Dai, Lixue Xia, Yunchao Wei, Yunhe Wang, Yu Wang
RankEvolve is an auto‑research framework that evolves generative ranking models by orchestrating multiple large‑language‑model coding agents through an Executable Operating Protocol (EOP). The system compiles a state machine that enforces phases, gates, branches, and loops, while a meta‑meta‑harness lets agents review and repair each other’s code. In budget‑matched experiments, heterogeneous composition of agents raised execution accuracy from 45.8 % to 62.5 % and reduced silent critical‑defect rates, achieving notable gains on the HSTU recommender and other benchmarks.
By Zheng Chen, Linfeng Liu, Hong Li, Hong Yan