arXiv AI By Hiroyasu Usami, Keisuke Hara, Ayato Tsuboi, Naohiko Matsuda

LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation

Read the original on arXiv AI →

arXiv:2606. 15610v1 Announce Type: cross Abstract: LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and difficult to reproduce.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels

The study investigates how the identity of a judge influences outcomes in large language model (LLM)-as-judge panels, using a fully crossed pairwise design across four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) with 9,312 judgments. A common per‑family statistic was found to be strongly confounded with candidate quality, prompting the authors to develop a corrected estimator that isolates judge effects while holding candidate family constant. The corrected analysis reveals a consistent positive same‑family lift (3.4–8.4 percentage points) across all families, with a global effect size of 0.067 (95 % CI [0.053, 0.084]) and a permutation p = 0.0002, and demonstrates that judge‑side likelihood and panel composition significantly influence outcomes.

By David Ababio Awuni, Luke E. K. Achenie, Benjamin Tei Partey, Elvis Gyasi Owusu, Nii-Nai Derrick Sowah
arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand
arXiv AI
Aug 26

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.

By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong