Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.
arXiv:2608.29675v1 Announce Type: cross Abstract: Repository exploration is a distinct and costly stage of coding-agent pipelines: before generating a patch, an agent must identify which repository f...
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks...
arXiv:2609.37143v1 Announce Type: cross Abstract: Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks wi...
arXiv:2606.21804v2 Announce Type: replace-cross Abstract: Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding age...
E2E-SWE is a benchmark that tests large language models’ ability to create complete, functional software repositories from scratch. It includes 186 tasks across 11 programming languages, each requiring an agent to build an installable project based solely on a natural‑language specification and an empty workspace, while passing a hidden test suite. The benchmark was crafted by software engineers and LLMs, then refined through iterative verification by autonomous agents to ensure clarity and solvability.