The study investigates how the identity of a judge influences outcomes in large language model (LLM)-as-judge panels, using a fully crossed pairwise design across four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) with 9,312 judgments. A common per‑family statistic was found to be strongly confounded with candidate quality, prompting the authors to develop a corrected estimator that isolates judge effects while holding candidate family constant. The corrected analysis reveals a consistent positive same‑family lift (3.4–8.4 percentage points) across all families, with a global effect size of 0.067 (95 % CI [0.053, 0.084]) and a permutation p = 0.0002, and demonstrates that judge‑side likelihood and panel composition significantly influence outcomes.
By David Ababio Awuni, Luke E. K. Achenie, Benjamin Tei Partey, Elvis Gyasi Owusu, Nii-Nai Derrick Sowah
The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.
By Atul Anand
arXiv:2606. 15474v1 Announce Type: new Abstract: Continuous evaluation of LLM products relies on a strong LLM judge treated as ground truth: a cheap monitor scores every interaction and a team is paged when the score drifts down.
By Yitao Li
arXiv:2609.22599v1 Announce Type: new
Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to...
By Jackson Hassell, Farima Fatahi Bayat, Pouya Pezeshkpour, Estevam Hruschka
The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.
By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
arXiv:2607. 08535v1 Announce Type: cross Abstract: An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed.
By Zongyou Yang, Yinghan Hou, Xiaokun Yang