arXiv AI By Frances Liu, Manny Silva, Paige Calvert, Ayu Adiati, Sarah Sanders

DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

Hugging Face Trending Papers
Aug 20

From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation

The study investigates how autonomous coding agents interact with technical documentation, analyzing 557 coding sessions and 33,097 pull requests. Findings reveal that agents primarily engage with agent-facing artefacts, show weak links between documentation consultation and code editing, lack explicit validation sequences, and tend to consult documentation after code changes. The authors propose a two‑lobed cycle model of agent‑documentation interaction and challenge assumptions about actionability and verifiability of agent‑friendly documentation.

arXiv AI
4d ago

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.

By Md Shohel Arman, Igor Molybog
arXiv AI
Sep 25

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

SWE-Prometheus is a new benchmark that evaluates large language model coding agents on the broader task of improving repository engineering governance, rather than just fixing individual issues. It presents fixed snapshots with open-ended objectives, requiring agents to identify risks, prioritize interventions, and verify changes across six governance dimensions. The benchmark includes 60 repositories and reports metrics such as Normalized Governance Improvement and behavior‑breakage rates, providing a nuanced view of how different models affect governance artifacts and execution‑backed improvements.

By Jiajun Wu, Leixin Sun, Zihan Tan, Yitao Liu, Shuo Li, Jiaru Qian, Shanghaoran Quan, Chuangxin Zhao, Yangxu Liao, Yang Liu, Bin Chong, Guancheng Wan