arXiv Computation and Language

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

The paper evaluates a production text‑to‑SQL pipeline that uses an LLM as a judge, finding that the deployed gpt‑4o‑mini judge agrees with human annotators only weakly (Cohen’s kappa 0.04 on a disagreement‑enriched set and 0.42 on a random spot‑check). The authors identify a specific failure mode, GRADE‑HALLUCINATION, responsible for most over‑flags, and demonstrate that a self‑hosted Qwen3.6‑27B model achieves substantially higher agreement (kappa 0.72) at a much lower cost. They also show that ensembling judges does not improve performance, and that their audit method flags a significant portion of out‑of‑domain SQLs as potential issues.

arXiv AI
Sep 2

Commit-first LLM judging inherits the judge's own errors

The paper investigates whether widely used evaluation frameworks for large language models (LLMs) implement a defense called commit‑first judging, which requires a judge to solve a task itself before accepting a candidate answer. Across 24 configurations in eight popular frameworks, none use the full commit‑first method; nine use a weaker variant that is ineffective. In controlled experiments, the weaker variant allowed systems to game the judge, while the full commit‑first approach eliminated this vulnerability but sometimes worsened evaluation when the judge’s own answer was incorrect.

By Idil Gozel
arXiv AI
Aug 26

More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight

The paper introduces the twin‑prefix framework to evaluate how the size of the verification unit—i.e., how many actions a pre‑execution LLM monitor reviews in one call—affects its performance. By pairing each gold plan with a twin that differs by a single write and injecting a controlled error, the authors isolate the impact of review length on catch rates and false rejections. Their findings show that longer review windows increase rejection rates but do not improve discrimination, with the highest informedness occurring at one or two actions across all judges and domains.

By Yuchen Han, Cheng Yan, Wuyang Zhang