arXiv:2609.07586v1 Announce Type: new
Abstract: Software repositories contain vast amounts of data on code contributions, bug reports, and project activities, yet this information remains challenging...
By Muhammad Jawad Chowdhury, Md. Sakib Khan
The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.
By Md Shohel Arman, Igor Molybog
arXiv:2604. 23816v2 Announce Type: replace-cross Abstract: Software documentation frequently becomes outdated or fails to exist entirely, yet developers need focused views of their codebase to understand complex systems.
By Oleg Baryshnikov, Anton M. Alekseev, Sergey I. Nikolenko
arXiv:2509.21891v3 Announce Type: replace-cross
Abstract: Fine-tuning large language models for code editing has typically relied on mining commits and pull requests. The working hypothesis has been...
By Yangtian Zi, Zixuan Wu, Aleksander Boruch-Gruszecki, Jonathan Bell, Arjun Guha
The paper reports on building a research-software catalog using a coding agent, starting from a three‑day hackathon prototype and moving to public deployment. It details the engineering work needed—adversarial review, data‑quality checks, browser validation, and publication safeguards—to ensure reliable operation, noting that silent failures were more problematic than crashes. The authors then examine applying these lessons to a larger, human‑curated portal (MateriApps) that combines curated metadata, external documentation, vector search, and local language‑model generation, finding that explicit validation, monitoring, and repeated review remain essential for AI‑assisted software portals.
By Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada
The study investigates how autonomous coding agents interact with technical documentation, analyzing 557 coding sessions and 33,097 pull requests. Findings reveal that agents primarily engage with agent-facing artefacts, show weak links between documentation consultation and code editing, lack explicit validation sequences, and tend to consult documentation after code changes. The authors propose a two‑lobed cycle model of agent‑documentation interaction and challenge assumptions about actionability and verifiability of agent‑friendly documentation.
arXiv:2608. 20195v1 Announce Type: cross Abstract: Technical documentation is written for human developers, but an increasing share of software changes is now authored by autonomous coding agents.
By Zhijun Gao, Jing Chen
arXiv:2606. 09090v1 Announce Type: cross Abstract: Developers increasingly provide AI coding assistants with persistent context through configuration files such as CLAUDE.
By Christoph Treude, Sebastian Baltes
arXiv:2408. 03910v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories.
By Xiangyan Liu, Bo Lan, Zhiyuan Hu, Yang Liu, Zhicheng Zhang, Fei Wang, Michael Shieh, Wenmeng Zhou
arXiv:2605.14563v3 Announce Type: replace-cross
Abstract: Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and...
By Suyoung Bae, Jaehoon Lee, Changkyu Choi, YunSeok Choi, Jee-Hyong Lee
arXiv:2608. 13662v1 Announce Type: new Abstract: Coding agents have become the primary means of generating new code in many software projects, and the resulting velocity of changes makes keeping track of the reasons behind those changes challenging.
By James Adam
The paper introduces CodeGraph, an open‑taxonomy knowledge graph that semantically annotates source code by extracting entities such as algorithms, paradigms, design patterns, and application domains from millions of files. Using a specialized large language model and a three‑stage Wikidata linking process, the authors ground these entities in Wikidata and construct a graph with about 158 million nodes and 1 billion typed edges across 14 programming languages. A quality‑assurance protocol combining human evaluation and an LLM‑as‑a‑judge filter quantifies annotation precision.
By Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina