arXiv AI

Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

The paper introduces a method for co-evolving inspectable verifiers alongside self-improving agents. By synthesizing verifiers from clustered failures and selecting them based on agreement with a reference set, the authors demonstrate improved held‑out agreement on MBPP+ and outperform a bare LLM judge. They also show that removing anchor guards collapses the verifier into a vacuous grader, yet the collapsed verifier still trains skills effectively, indicating that downstream task scores cannot certify a self‑evolved verifier.

arXiv AI
Sep 3

LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

The paper argues that using a large language model (LLM) as the sole judge in self‑improving agent pipelines is problematic, as the judge can be biased or manipulated, leading to false confidence in system performance. The authors propose a new framework, PROCTOR, which replaces the oracle judge with a deterministic, teacher‑student loop that enforces guardrails such as sandboxing, role separation, and acceptance checks to prevent cheating and ensure reliable evaluation. Experiments across contract analysis, compliance review, and code quality demonstrate that PROCTOR mitigates eleven identified failure modes that previously allowed agents to achieve perfect scores while hiding significant capability gaps.

By Vansh Wahi
arXiv AI
Jun 30

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

arXiv:2606. 29713v1 Announce Type: cross Abstract: Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers emit only opaque binary labels, leaving agents unable to self-correct and operators unable to audit.

By Aojie Yuan, Yi Nian, Haiyue Zhang, Zijian Su, Yue Zhao
arXiv AI
6d ago

VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

The paper introduces VERSE, a Verified Self‑Evolving optimizer that enhances LLM agent harnesses by allowing the optimizer to test edits, replay failures, and perturb steps while tracking fixes and regressions. VERSE builds its own tools for failure analysis, verification, training audits, and workflow control, and uses this feedback to revise the harness’s prompts, skills, tools, hooks, and notes without changing model weights. In experiments across five executors and multiple languages, VERSE improves all evaluated harness optimizers, achieving higher accuracy on held‑out and out‑of‑distribution tasks compared to the strongest baselines.

By Zekai Wang, Yingqiang Ge, Zekun Wang, Hai Wang, Yuhui Xu, Joshua Frandsen, Shancong Fu, Ashia C. Wilson, Chandan K. Reddy
arXiv AI
Aug 7

When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents

arXiv:2608. 05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of improving it.

By Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, Ning Zheng
arXiv Computation and Language
2d ago

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

The paper introduces TRACE, a protocol designed to diagnose whether changes in verifier scores for large language model agents reflect actual changes in agent behavior or merely alterations in the evaluation process. TRACE works by applying targeted changes to evaluation components, running paired experiments, and re‑scoring unchanged trajectories to isolate the source of score variation. Experiments on synthetic tasks and public benchmark tasks demonstrate that seemingly significant score shifts can often be attributed to evaluation artifacts, while TRACE can also detect genuine behavioral changes.

By Radhika Gaonkar