arXiv AI By David Ababio Awuni, Luke E. K. Achenie, Benjamin Tei Partey, Elvis Gyasi Owusu, Nii-Nai Derrick Sowah

Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels

Read the original on arXiv AI →

The study investigates how the identity of a judge influences outcomes in large language model (LLM)-as-judge panels, using a fully crossed pairwise design across four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) with 9,312 judgments. A common per‑family statistic was found to be strongly confounded with candidate quality, prompting the authors to develop a corrected estimator that isolates judge effects while holding candidate family constant. The corrected analysis reveals a consistent positive same‑family lift (3.4–8.4 percentage points) across all families, with a global effect size of 0.067 (95 % CI [0.053, 0.084]) and a permutation p = 0.0002, and demonstrates that judge‑side likelihood and panel composition significantly influence outcomes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.

By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand
arXiv AI
2d ago

The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

The paper demonstrates that reading a large language model (LLM) judge’s verdict from the logits of its first generated token—an approach used in constrained decoding and likelihood‑scoring evaluation—introduces a significant distortion in position bias. Because judges do not always start with a verdict token (12–49% of cases for Qwen3 judges and <3% for Llama‑3.1‑8B and Phi‑3.5‑mini), this readout often returns the first response rather than a true judgment, inflating position bias by up to 42 points while barely affecting judge accuracy. The authors recommend reporting the frequency with which a judge leads with a verdict token to provide a more accurate assessment of position bias.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli