arXiv AI

When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess

The paper investigates when a language‑model judge can truly ground its verdicts in code correctness. It shows that current multi‑agent verification methods rely on evidence that is both independent of the answer and distinct between candidates—conditions that fail in code judging. By analyzing two label‑free measurements from the judge’s logs, the authors demonstrate that gating on one measurement allows the system to decline uncertain comparisons, improving accuracy from 20.7% to 36.9% while still answering half of all cases.

arXiv Computation and Language
Sep 21

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

JudgeSense is a benchmark comprising 880 items from human‑labelled corpora, each presented under two differently worded instructions that ask the same question. The study evaluates 25 judges from six providers across four tasks, measuring how rewording affects agreement with the judge’s own verdicts. Results show that rewording reduces agreement on all tasks, with significant effects on two, and that stability varies across tasks and is not predicted by parameter count.

By Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang
arXiv AI
Sep 2

Commit-first LLM judging inherits the judge's own errors

The paper investigates whether widely used evaluation frameworks for large language models (LLMs) implement a defense called commit‑first judging, which requires a judge to solve a task itself before accepting a candidate answer. Across 24 configurations in eight popular frameworks, none use the full commit‑first method; nine use a weaker variant that is ineffective. In controlled experiments, the weaker variant allowed systems to game the judge, while the full commit‑first approach eliminated this vulnerability but sometimes worsened evaluation when the judge’s own answer was incorrect.

By Idil Gozel
arXiv Machine Learning
1d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
Aug 26

Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks

The paper introduces a dual‑judge evaluation protocol for vision‑language models in legally grounded tasks, pairing a 0‑10 quality judge with a strict binary semantic‑equivalence judge. Using a controlled UK traffic‑sign interpretation task, the authors analyze 4,680 evaluations across visibility and occlusion conditions, finding moderate association between judges and an asymmetric Type II error pattern that is most pronounced under heavy occlusion. The protocol requires only one additional LLM call and reveals quality‑trustworthiness signals that single‑judge methods miss.

By Su Myat Noe, Ha Thanh Nguyen, May Myo Zin, Ken Satoh
arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin
arXiv AI
Aug 24

Structure for Reading, Prose for Writing: Asymmetric Structural Conditioning in Multi-Agent Document Authoring

The paper reports on a deployed multi‑agent tender‑response system that uses an open‑weights language model under sovereignty constraints. In a blind comparison, the system’s answers were judged at least as good as human‑written bids in 40 of 55 sections, with only a few gaps attributable to missing knowledge rather than writing quality. The study also demonstrates an asymmetry in conditioning: while structural markup improves reading tasks, converting instruction material from prose to nested XML degrades answer quality, and naming forbidden constructions concentrates defects.

By Cheng Yu, Nikhil Mathew, Zhengjie Wang