arXiv:2607. 17641v1 Announce Type: new Abstract: Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use.
By Yitao Wu, Si Shen, Rui Yang, Hong Peng, Bin Hu
The paper introduces LLM-as-an-Improver, a method that uses verification feedback to enhance the candidate set in verifier-based selection. It proposes Verify–Repair–Reselect (VRR), which keeps the initial winner, generates three complementary alternatives (repaired versions of the winner and runner‑up, and a new approach), filters invalid or duplicate candidates, and then reselects the final answer. Experiments on code‑generation and reasoning benchmarks show that VRR outperforms fixed‑pool selection and can recover correct solutions even when the initial pool is entirely wrong.
By Akiyoshi Tomihari, Yuma Ichikawa
arXiv:2607. 14167v1 Announce Type: cross Abstract: LLM agents often retry after external validation rejects a candidate, but the interface between validation and the next model call remains underspecified.
By Jaideep Ray, Ankit Goyal
ContractRL introduces a contract-constrained sequential repair protocol for structured tool calls, modeling verifier-guided JSON repair as a bounded decision process. The policy observes candidate data, verifier feedback, JSON pointers, repair history, and budget, using a contract-derived action mask to filter invalid operations before a deterministic validator applies changes. Compared to Patch‑SFT and full regeneration, ContractRL achieves higher semantic success (0.9362 vs. 0.9076 and 0.9148) while generating fewer tokens (34.4 vs. 44.9 and 137.2), and policy optimization further improves success rates.
By Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao
arXiv:2606. 19808v1 Announce Type: new Abstract: Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-correct answers, or introduce harmful answer changes.
By Sajib Acharjee Dip, Dawei Zhou, Liqing Zhang
arXiv:2607. 24604v1 Announce Type: cross Abstract: Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee.
By Xueping Gao, Jianwei Yang, Qiang Yang