arXiv:2607. 17641v1 Announce Type: new Abstract: Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use.
By Yitao Wu, Si Shen, Rui Yang, Hong Peng, Bin Hu
Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use. When both the verifier and the repairer are noisy, repair can damage already-correct plans, and reported acceptance keeps rising while true validity falls, so existing methods lack a principled basis for deciding when repair should stop.
The Agent Error Dataset (AED) presents 50,228 error–diagnosis pairs collected from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text‑based agent systems. A five‑stage Agentic Error‑to‑Training (AET) pipeline generates diagnoses and proposed corrections, verifies them against recorded evidence, and creates separate training views for diagnosis and actor recovery. Experiments show that first‑proposal corrections improve verifier pass rates from 18.4% to 51.1%, and fine‑tuning with full‑diagnosis data raises Qwen3‑8B’s exact‑step agreement from 47.2% to 63.6% on a holdout set.
By Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li, Heng Ji
arXiv:2609.15684v1 Announce Type: new
Abstract: Language agents increasingly rely on reusable skills, but post-failure repair is often handled by opaque one-shot reflection: a model generates a skill...
By Mengyi Deng, Xin Li, Duyi Pan, Zilin Wang, Zhiwei Li, Zhijiang Guo, Wei Wang
The paper introduces a fixed‑budget revision protocol that uses deterministic verifiers to expose all remaining violations across exact‑length, lexical, and compositional constraints, thereby isolating model‑side revision behavior. Experiments on 19 open‑ and closed‑source LLMs show wide variability in controller‑level success, with some models achieving up to 99.8% success while others remain below 20%. Controlled studies reveal that post‑training and scale affect model responses to exact feedback, but do not consistently improve exact correction, and that recurrence of earlier outputs is linked to lower recoverability.
By Haitong Jiang, Chunlin Liu, Yile Wang, Yuhong Feng
arXiv:2607. 28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect.
By Xiaonan Xu, Wenjing Wu