The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.
By Md Shohel Arman, Igor Molybog
arXiv:2606. 09090v1 Announce Type: cross Abstract: Developers increasingly provide AI coding assistants with persistent context through configuration files such as CLAUDE.
By Christoph Treude, Sebastian Baltes
arXiv:2609. 31176v1 Announce Type: new Abstract: Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue.
By Yunxiang Wei, Zhenyu Lei, Jundong Li
arXiv:2607. 01425v1 Announce Type: new Abstract: Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge.
By Yongjian Tang, Ezgi Sarikayak, Doruk Tuncel, Jie M. Zhang, Thomas Runkler
arXiv:2507. 16395v3 Announce Type: replace Abstract: Atomic commits, which address a single development concern, are a best practice in software development.
By Bo Hou, Xin Tan, Kai Zheng, Fang Liu, Yinghao Zhu, Li Zhang
arXiv:2607. 08691v1 Announce Type: cross Abstract: Repository-level code generation requires implementing target functions while accounting for complex cross-file dependencies and project-specific conventions.
By QiHong Chen, Aaron Imani, Iftekhar Ahmed
arXiv:2606. 28379v1 Announce Type: cross Abstract: We introduce LEDGER to tackle the novel context engineering challenge of agentic document editing, where localized edits to long, structured documents must be applied efficiently without breaking cross-references or semantic consistency.
By Mike Hang Wang, Utkarsh Garg, Reza Davari, Huitian Jiao, Hao Cheng, Baolin Peng, Tao Ge, Si-Qing Chen
arXiv:2602.11988v3 Announce Type: replace-cross
Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although thi...
By Thibaud Gloaguen, Niels M\"undler-Sasahara, Mark Niklas M\"uller, Veselin Raychev, Martin Vechev
arXiv:2608.29310v1 Announce Type: cross
Abstract: Organizations often develop and maintain portfolios of related applications: independently deployable codebases that share substantial domain logic,...
By Daegyu Sung, Yukyeong Lee, Geon Park, Yumin Choi, Sung Ju Hwang
arXiv:2608. 10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into explicit execution structures.
By Shuo Hao, You Lu, Bihuan Chen, Xin Peng
arXiv:2609.39765v1 Announce Type: new
Abstract: Agent memory faces heterogeneous access needs: a single-hop question may require one piece of evidence, whereas a multi-hop question must combine evide...
By Xiaoqiang Wang, Bang Liu
E2E-SWE is a benchmark that tests large language models’ ability to create complete, functional software repositories from scratch. It includes 186 tasks across 11 programming languages, each requiring an agent to build an installable project based solely on a natural‑language specification and an empty workspace, while passing a hidden test suite. The benchmark was crafted by software engineers and LLMs, then refined through iterative verification by autonomous agents to ensure clarity and solvability.
By Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou