arXiv AI

Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases

arXiv:2607. 01425v1 Announce Type: new Abstract: Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge.

arXiv AI
6d ago

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.

By Md Shohel Arman, Igor Molybog
arXiv AI
Sep 2

WiseSpec: Requirements-Driven Agents for Code Generation

WiseSpec is a requirements‑driven agent framework designed to improve repository‑level code generation. It automatically builds structured, information‑rich requirements, evaluates their quality via execution‑based tests, and iteratively refines them to better guide code generation. Experiments show WiseSpec outperforms all baselines, achieving an average 13.17% improvement in %Resolved.

By Zhao Tian
arXiv AI
6d ago

BabelCoder: Agentic Code Translation with Specification Alignment

BabelCoder is an agentic framework for automatic code translation that splits the task into specialized agents for translation, testing, and refinement. Each agent focuses on a specific aspect—generating code, validating correctness, or repairing errors—allowing collaborative improvement of translation quality. Evaluated on four benchmark datasets, BabelCoder outperforms four state‑of‑the‑art baselines, achieving an average accuracy of 94.16% and surpassing existing methods in 94% of cases.

By Fazle Rabbi, Soumit Kanti Saha, Tri Minh Triet Pham, Song Wang, Jinqiu Yang
arXiv AI
Sep 4

Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection

Dude is a Dual-Detection Multi-Agent System designed to detect discrepancies between research papers and their accompanying code. It addresses limitations of single-agent LLM approaches by aligning granularity between paper and code through negotiation and a two-stage salience-filtering mechanism, reducing false positives. Experiments on real-world datasets show that Dude improves recall and precision by up to 22.8% and boosts the F1 score by up to 18.7% over baseline methods.

By Weijie Liu, Running Zhao, Wenhao Yuan, Jinfeng Xu, Zhanfeng Xu, Xiaoxi Zhang, Edith Cheuk-Han Ngai
arXiv Machine Learning
Aug 6

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories

arXiv:2604. 07341v2 Announce Type: replace-cross Abstract: Most repository-level code translation and validation techniques have been evaluated on a single source-target programming language (PL) pair, owing to the complex engineering effort required to adapt new PL pairs.

By Ali Reza Ibrahimzada, Brandon Paulsen, Daniel Kroening, Reyhaneh Jabbarvand
arXiv AI
Aug 5

AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection

arXiv:2601. 19138v2 Announce Type: replace-cross Abstract: Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security assessment is deferred to later stages, delaying feedback and increasing remediation costs.

By Wachiraphan Charoenwet, Kla Tantithamthavorn, Patanamon Thongtanunam, Hong Yi Lin, Minwoo Jeong, Ming Wu