arXiv AI

Cherry-pick Override: LLM Judges Under-use the Non-Directional Verdicts Their Contract Authorizes

arXiv AI
Sep 2

Commit-first LLM judging inherits the judge's own errors

The paper investigates whether widely used evaluation frameworks for large language models (LLMs) implement a defense called commit‑first judging, which requires a judge to solve a task itself before accepting a candidate answer. Across 24 configurations in eight popular frameworks, none use the full commit‑first method; nine use a weaker variant that is ineffective. In controlled experiments, the weaker variant allowed systems to game the judge, while the full commit‑first approach eliminated this vulnerability but sometimes worsened evaluation when the judge’s own answer was incorrect.

By Idil Gozel
arXiv AI
Aug 24

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

JuryProbe is an empirical diagnostic tool designed to assess consensus risk in panels of reference‑free large language model judges used for factuality verification. It estimates risk by measuring false‑negative correlations and false‑consensus lift from a labeled calibration probe, and routes high‑risk majority decisions to judges with trusted references. The approach was validated on FEVER corruptions, showing that flagged decisions can be grounded without additional reference acquisition in most cases, while reducing false accepts by about 0.4% and avoiding 28% of reference acquisitions.

By Tianxin Zhou, Ruixi Lin
arXiv AI
Aug 26

More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight

The paper introduces the twin‑prefix framework to evaluate how the size of the verification unit—i.e., how many actions a pre‑execution LLM monitor reviews in one call—affects its performance. By pairing each gold plan with a twin that differs by a single write and injecting a controlled error, the authors isolate the impact of review length on catch rates and false rejections. Their findings show that longer review windows increase rejection rates but do not improve discrimination, with the highest informedness occurring at one or two actions across all judges and domains.

By Yuchen Han, Cheng Yan, Wuyang Zhang
arXiv Machine Learning
1d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv Computation and Language
3d ago

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

The paper evaluates a production text‑to‑SQL pipeline that uses an LLM as a judge, finding that the deployed gpt‑4o‑mini judge agrees with human annotators only weakly (Cohen’s kappa 0.04 on a disagreement‑enriched set and 0.42 on a random spot‑check). The authors identify a specific failure mode, GRADE‑HALLUCINATION, responsible for most over‑flags, and demonstrate that a self‑hosted Qwen3.6‑27B model achieves substantially higher agreement (kappa 0.72) at a much lower cost. They also show that ensembling judges does not improve performance, and that their audit method flags a significant portion of out‑of‑domain SQLs as potential issues.

By Haowei Liu, Hsin-Tai Wu, Yi Fang