arXiv AI By Yongjian Tang, Ezgi Sarikayak, Doruk Tuncel, Jie M. Zhang, Thomas Runkler

Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases

Read the original on arXiv AI →

arXiv:2607. 01425v1 Announce Type: new Abstract: Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.

By Md Shohel Arman, Igor Molybog
arXiv AI
Sep 2

WiseSpec: Requirements-Driven Agents for Code Generation

WiseSpec is a requirements‑driven agent framework designed to improve repository‑level code generation. It automatically builds structured, information‑rich requirements, evaluates their quality via execution‑based tests, and iteratively refines them to better guide code generation. Experiments show WiseSpec outperforms all baselines, achieving an average 13.17% improvement in %Resolved.

By Zhao Tian
arXiv AI
6d ago

BabelCoder: Agentic Code Translation with Specification Alignment

BabelCoder is an agentic framework for automatic code translation that splits the task into specialized agents for translation, testing, and refinement. Each agent focuses on a specific aspect—generating code, validating correctness, or repairing errors—allowing collaborative improvement of translation quality. Evaluated on four benchmark datasets, BabelCoder outperforms four state‑of‑the‑art baselines, achieving an average accuracy of 94.16% and surpassing existing methods in 94% of cases.

By Fazle Rabbi, Soumit Kanti Saha, Tri Minh Triet Pham, Song Wang, Jinqiu Yang