arXiv AI

TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning

arXiv:2607. 26307v1 Announce Type: new Abstract: Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible.

arXiv Computation and Language
Aug 27

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

The paper argues that answer accuracy alone is insufficient for evaluating large language model (LLM) data agents, especially in structured-data tasks where a correct answer can be produced by an invalid trace. It introduces Trace Integrity as a reliability criterion that ensures the computation behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. The authors operationalize this concept with execution contracts and present the CAIT (Correct Answer / Invalid Trace) Rate to quantify how often answer-only evaluations mistakenly reward unsupported outputs, demonstrating that accuracy, trace validity, and silent-failure risk are distinct signals.

By Srimonti Dutta, Akshata Kishore Moharir
arXiv AI
Sep 15

Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program Repair

The paper introduces THEMIS, a stage-aware repair workflow that externalizes the requirement-to-repair process by generating semantic interpretations, a runtime requirement-code graph, graph-derived developer guidance, retained repair rationale and patches, and post-edit audit records. A retrospective audit of 300 SWE-bench Lite cases shows that these artifacts enable cross-stage inspection, with a complete developer rationale available for 288 cases and 214 cases retaining a full audited field set. The retained records also allow systematic measurement of cross-stage correspondence, revealing high recurrence of target symbols across rationales and patches, and a preliminary improvement in resolving cases compared to a direct same-input condition.

By Zewen Tao, Shin-nosuke Ishikawa
arXiv AI
Aug 20

LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents

LEDGER is a tracing and review system for large language model agents that constructs layered trace graphs from observed sessions. It groups raw trace records into Evidence Nodes and Workflow Nodes, anchors artifacts as evidence, and adds typed semantic edges linking claims to supporting actions, artifacts, and checks. The resulting traces reveal workflow decisions, artifact lineage, repair steps, validation coverage, and claim‑support paths for evidence‑centered audit.

By Daehong Kim, Haichao Miao, Shusen Liu
arXiv AI
Aug 11

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

arXiv:2608. 09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.

By Xin Zhou, Chun Yong Chong, Kisub Kim, Yun Peng, Rui Shu, Zihan Wu, Xu Han, Guowen Yuan, Zeyang Zhuang, Jounghoon Kim, Jeongjin Ju, Seongmin Ju, Taein Yoon, David Lo
arXiv AI
4d ago

StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

StateTape introduces a new framework for long‑horizon coding agents that rewrites the agent’s context as the code repository changes, rather than letting the context grow with every observation. It models the repository as a symbol‑level code graph, using a tape to mark symbols altered by each write and a manager model to resolve stale records. The authors provide theoretical analysis, a new benchmark called TraceBench, and empirical results showing higher resolve rates across six agents and three edit‑heavy benchmarks with minimal computational overhead.

By Ziyang Yu, Liang Zhao, Bowen Zhu, Hasibul Haque