arXiv AI

Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents

arXiv:2607. 17641v1 Announce Type: new Abstract: Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use.

Hugging Face Trending Papers
Jul 20

Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents

Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use. When both the verifier and the repairer are noisy, repair can damage already-correct plans, and reported acceptance keeps rising while true validity falls, so existing methods lack a principled basis for deciding when repair should stop.

arXiv AI
Sep 18

LLM-as-an-Improver: Turning Verification into Better Candidates

The paper introduces LLM-as-an-Improver, a method that uses verification feedback to enhance the candidate set in verifier-based selection. It proposes Verify–Repair–Reselect (VRR), which keeps the initial winner, generates three complementary alternatives (repaired versions of the winner and runner‑up, and a new approach), filters invalid or duplicate candidates, and then reselects the final answer. Experiments on code‑generation and reasoning benchmarks show that VRR outperforms fixed‑pool selection and can recover correct solutions even when the initial pool is entirely wrong.

By Akiyoshi Tomihari, Yuma Ichikawa
arXiv AI
2d ago

ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair

ContractRL introduces a contract-constrained sequential repair protocol for structured tool calls, modeling verifier-guided JSON repair as a bounded decision process. The policy observes candidate data, verifier feedback, JSON pointers, repair history, and budget, using a contract-derived action mask to filter invalid operations before a deterministic validator applies changes. Compared to Patch‑SFT and full regeneration, ContractRL achieves higher semantic success (0.9362 vs. 0.9076 and 0.9148) while generating fewer tokens (34.4 vs. 44.9 and 137.2), and policy optimization further improves success rates.

By Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao
arXiv AI
Jul 23

Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes

arXiv:2607. 19843v1 Announce Type: cross Abstract: Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly from bug reports remains underconstrained.

By Yuhao Tan, Zhibang Yang, Fangkai Yang, Yuan Yao, Yu Kang, Lu Wang, Pu Zhao, Xin Zhang, Xiaoxing Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
arXiv Machine Learning
Jul 1

Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

arXiv:2606. 31630v1 Announce Type: new Abstract: Language models increasingly write probabilistic programs (in NumPyro, Stan, or Pyro), but a program that compiles, runs, and passes every unit test can still be \emph{statistically} wrong -- a Gaussian likelihood for heavy-tailed data, a Poisson for over-dispersed counts, an invalid prior support, or a pathological parameterization.

By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv Computation and Language
Aug 31

Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models

The study investigates how language models equipped with tools can still produce unsupported final claims, even when a single tool call could resolve the uncertainty. It defines two metrics—occurrence (how often unsupported claims arise) and conditional repair (how often they are fixed when evidence is provided). Experiments on Qwen3-32B and Gemma 4 show that providing the missing evidence consistently repairs all unsupported claims in the Qwen3-32B setup, while the Gemma 4 model never produced unsupported claims under the tested conditions.

By Justin Bronder