arXiv Computation and Language

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

The paper introduces TRACE, a protocol designed to diagnose whether changes in verifier scores for large language model agents reflect actual changes in agent behavior or merely alterations in the evaluation process. TRACE works by applying targeted changes to evaluation components, running paired experiments, and re‑scoring unchanged trajectories to isolate the source of score variation. Experiments on synthetic tasks and public benchmark tasks demonstrate that seemingly significant score shifts can often be attributed to evaluation artifacts, while TRACE can also detect genuine behavioral changes.

arXiv Machine Learning
Sep 30

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
Sep 15

Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs

The paper introduces a new evaluation method called "same-input rerun" to assess the consistency of clinical language‑model agents across repeated runs. By replaying 1,000 MedAgentBench tasks with identical inputs, the authors find that action‑level outputs—such as test orders, medication requests, and referrals—vary significantly, even when benchmark scores remain unchanged. The study demonstrates that current benchmarks, which typically evaluate only a single run per task, can miss substantial behavioral divergence.

By Rohith Reddy Bellibatlu, Manpreet Singh, Zhoutian Han, Wenbin Zhang
arXiv Machine Learning
Sep 25

Don't Read the Log: Execution Traces Contaminate Verifiers in Video-Generation Agents

The paper investigates how providing execution traces to multimodal judges in agentic video‑generation systems can bias their verdicts. On a benchmark of 109 two‑event clips, traces that falsely report successful tool calls cause large‑language‑model judges to incorrectly accept 78–90 % of failures, while contradictory traces lead to 100 % rejection of correct clips. The effect persists even when judges are instructed to consider only the video frames, indicating that the vulnerability stems from the judges’ learned trust in tool logs rather than the visual content itself.

By Jian Xu
arXiv AI
Sep 2

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

The paper "trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories" examines the limitations of outcome-only evaluation for large language model agents. Using a deterministic tool‑using support‑desk environment with a scripted oracle policy and a fault injector, the authors compare five different judging approaches—programmatic rules, outcome‑only, step‑rubric at two model sizes, and a self‑consistency ensemble—on metrics such as detection, step localisation, fault typing, calibration, and cost across 400 trajectories. The study finds that outcome‑only judges miss many silent faults and generate false positives, while step‑rubric judges achieve higher recall with no false alarms but at greater cost, and that none of the judges read the final reply, allowing fabricated promises to evade detection. "whyItMatters":"The findings highlight that current production‑default outcome‑only evaluations can overlook critical failures in agent behavior, underscoring the need for more nuanced, step‑level judging methods to ensure reliable LLM agent performance."

By Hadi Mohammadi
arXiv AI
2d ago

Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

The paper introduces a method for co-evolving inspectable verifiers alongside self-improving agents. By synthesizing verifiers from clustered failures and selecting them based on agreement with a reference set, the authors demonstrate improved held‑out agreement on MBPP+ and outperform a bare LLM judge. They also show that removing anchor guards collapses the verifier into a vacuous grader, yet the collapsed verifier still trains skills effectively, indicating that downstream task scores cannot certify a self‑evolved verifier.

By Xing Zhang, Guanghui Wang, Yanwei Cui, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He
arXiv AI
Sep 2

Commit-first LLM judging inherits the judge's own errors

The paper investigates whether widely used evaluation frameworks for large language models (LLMs) implement a defense called commit‑first judging, which requires a judge to solve a task itself before accepting a candidate answer. Across 24 configurations in eight popular frameworks, none use the full commit‑first method; nine use a weaker variant that is ineffective. In controlled experiments, the weaker variant allowed systems to game the judge, while the full commit‑first approach eliminated this vulnerability but sometimes worsened evaluation when the judge’s own answer was incorrect.

By Idil Gozel