arXiv AI

ARISE: A Repository-level Graph Representation and Toolset for Agentic Program Repair and Fault Localization

arXiv:2605. 03117v2 Announce Type: replace-cross Abstract: Automated program repair at repository scale requires an agent to locate a fault among thousands of files and synthesize a correct patch.

arXiv AI
Jul 3

BLAgent: Agentic RAG for File-Level Bug Localization

arXiv:2605. 17965v2 Announce Type: replace-cross Abstract: Bug localization remains a key bottleneck for large language model (LLM)-based software maintenance, where accurately identifying faulty code is essential for debugging, root cause analysis, triage, and automated program repair (APR).

By Md Afif Al Mamun, Gias Uddin
arXiv Machine Learning
Sep 14

ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents

ParaRecover is a new process-level benchmark designed to evaluate error localization and recovery in multi-turn parallel tool-use agents. It contains 10,626 instances across two difficulty levels, built on a taxonomy of 14 error types that cover planning dependencies, tool selection, and argument matching. The benchmark introduces the SDE rubric, which assesses structural integrity, diagnostic reasoning, and evolutionary strategy during agent execution, and demonstrates that it can guide improvements in agents’ reflective recovery capabilities.

By Bowen Guan, Zhentao Yin, Yanming Shen
arXiv AI
Sep 3

Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives

The paper introduces Tool Primitives, a design that replaces rigid API schemas with natural language interfaces for tool calling, enabling seamless inter-tool communication. It builds ToolFace, a repository of over 25,000 functions that LLMs can dynamically retrieve, and HEART, a harness engineering framework that orchestrates tool use with planning, routing, and verification. Experiments show HEART outperforms fine‑tuned models and leading commercial LLMs while cutting API costs by up to 85%.

By Haibo Jin, Suijin Wang, Xucheng Yu, Haojing Luo, Haohan Wang
arXiv AI
Aug 11

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

arXiv:2608. 09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.

By Xin Zhou, Chun Yong Chong, Kisub Kim, Yun Peng, Rui Shu, Zihan Wu, Xu Han, Guowen Yuan, Zeyang Zhuang, Jounghoon Kim, Jeongjin Ju, Seongmin Ju, Taein Yoon, David Lo
arXiv AI
Sep 15

Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program Repair

The paper introduces THEMIS, a stage-aware repair workflow that externalizes the requirement-to-repair process by generating semantic interpretations, a runtime requirement-code graph, graph-derived developer guidance, retained repair rationale and patches, and post-edit audit records. A retrospective audit of 300 SWE-bench Lite cases shows that these artifacts enable cross-stage inspection, with a complete developer rationale available for 288 cases and 214 cases retaining a full audited field set. The retained records also allow systematic measurement of cross-stage correspondence, revealing high recurrence of target symbols across rationales and patches, and a preliminary improvement in resolving cases compared to a direct same-input condition.

By Zewen Tao, Shin-nosuke Ishikawa