DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 20195v1 Announce Type: cross Abstract: Technical documentation is written for human developers, but an increasing share of software changes is now authored by autonomous coding agents.
The study investigates how autonomous coding agents interact with technical documentation, analyzing 557 coding sessions and 33,097 pull requests. Findings reveal that agents primarily engage with agent-facing artefacts, show weak links between documentation consultation and code editing, lack explicit validation sequences, and tend to consult documentation after code changes. The authors propose a two‑lobed cycle model of agent‑documentation interaction and challenge assumptions about actionability and verifiability of agent‑friendly documentation.
arXiv:2608. 10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents.
The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.
SWE-Prometheus is a new benchmark that evaluates large language model coding agents on the broader task of improving repository engineering governance, rather than just fixing individual issues. It presents fixed snapshots with open-ended objectives, requiring agents to identify risks, prioritize interventions, and verify changes across six governance dimensions. The benchmark includes 60 repositories and reports metrics such as Normalized Governance Improvement and behavior‑breakage rates, providing a nuanced view of how different models affect governance artifacts and execution‑backed improvements.
arXiv:2606. 13468v1 Announce Type: cross Abstract: AI coding agents are increasingly used to generate pull requests (PRs) that propose code fixes in software projects.