arXiv AI By Zewen Tao, Shin-nosuke Ishikawa

Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program Repair

Read the original on arXiv AI →

The paper introduces THEMIS, a stage-aware repair workflow that externalizes the requirement-to-repair process by generating semantic interpretations, a runtime requirement-code graph, graph-derived developer guidance, retained repair rationale and patches, and post-edit audit records. A retrospective audit of 300 SWE-bench Lite cases shows that these artifacts enable cross-stage inspection, with a complete developer rationale available for 288 cases and 214 cases retaining a full audited field set. The retained records also allow systematic measurement of cross-stage correspondence, revealing high recurrence of target symbols across rationales and patches, and a preliminary improvement in resolving cases compared to a direct same-input condition.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair

The paper investigates hallucination in large language model–based automated program repair (APR). It defines hallucination as producing patches or intermediate artifacts that are not grounded in available repair evidence, and analyzes it across final patches and intermediate tasks such as triggering test case identification, line coverage prediction, and additional test case generation. Experiments on 832 Defects4J bugs show that only 21.0%–55.9% of patches pass the developer test suite, with 72.7% of sampled repairs exhibiting hallucinations, often due to incorrect causal localization or repair strategies.

By Xuemeng Cai, Jiakun Liu, Linhan Yang, Wei Ma, Lingxiao Jiang
arXiv AI
Jul 31

TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning

arXiv:2607. 26307v1 Announce Type: new Abstract: Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible.

By Rwaida Alssadi, Muntaser Syed, Balaji Kasula, Lamine Deen, Majed Alotaibi, Mohammed Alghamdi, Tyler Ton, Ali Alqarni, Marius Silaghi