The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.
By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal
The study investigates how prior scores influence large language model (LLM) judgments in the LLM-as-a-Judge paradigm. By testing three prompt conditions—no metadata, revision framing, and anchored metadata containing prior scores—the authors find that prior scores systematically bias evaluations, shifting ratings toward those scores across 192,000 attempts. The bias also affects categorical decisions, blocking 48% of error corrections and flipping 10.18% of correct judgments, and is not mitigated by Chain-of-Thought or a warning, underscoring the need for careful context engineering.
By Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic
The paper argues that verbalized confidence—once viewed as overconfident and coarse—has become the preferred soft‑scoring method for LLM‑as‑a‑Judge on top‑tier proprietary models released after 2025. Experiments on SummEval, AggreFact, and HelpSteer2 across up to 18 LLMs show that log‑probabilities are no longer the best signal, and that adding an overconfidence advisory and self‑debate further improves calibration and robustness. The authors note that these enhancements incur little accuracy loss on post‑2025 models but do affect pre‑2025 ones, highlighting a compatibility shift in how confidence should be measured.
By Yu-Chung Hsiao
arXiv:2609.13773v1 Announce Type: new
Abstract: LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may...
By Ruichen Zheng, Yihe Wang, Fabrice Y Harel-Canada, Sara Khosravi, Zeynep Senahan Yildiz, Amit Sahai, Nanyun Peng
arXiv:2609.13824v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faste...
By Aakash Kumar Tiwari
arXiv:2608.20385v1 Announce Type: new
Abstract: Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and sensitive to ambiguity in checklist criteria. Al...
By Timo van der Kuil (Methodology and Statistics Utrecht University), Bruno Messina Coimbra (Methodology and Statistics Utrecht University), Mirjam van Zuiden (Clinical Psychology Utrecht University), Robert A. Bagheri (Methodology and Statistics Utrecht University), Rens van de Schoot (Methodology and Statistics Utrecht University), Klaas Dieleman (Methodology and Statistics Utrecht University), Berend Greijn (Methodology and Statistics Utrecht University), Stefan Houkes (Methodology and Statistics Utrecht University), Sebastiaan Rodenhuis (Methodology and Statistics Utrecht University), Elizabeth M. Grandfield (Methodology and Statistics Utrecht University)
JudgeSense is a benchmark comprising 880 items from human‑labelled corpora, each presented under two differently worded instructions that ask the same question. The study evaluates 25 judges from six providers across four tasks, measuring how rewording affects agreement with the judge’s own verdicts. Results show that rewording reduces agreement on all tasks, with significant effects on two, and that stability varies across tasks and is not predicted by parameter count.
By Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang
JudgeProfile is a framework that analyzes the subjectivity of large language model (LLM) judges by separating evaluation into perception—how judges compare responses on attributes such as clarity, correctness, and detail—and prioritization—how much each attribute influences the final decision. Using the curated SubjectiveSet dataset of 50,013 response pairs evaluated by 21 judges across 87 attributes, the study finds that judges often agree on attribute judgments even when their overall choices differ. By estimating and adjusting attribute weights, the authors improve agreement with reference labels from 66.48% to 71.97%, outperforming fine‑tuning and rubric prompting.
By Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, Yuanpeng Tu, Daniel Li, Junbiao Tang, Pengtao Xie, Zihao He
The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.
By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
arXiv:2606. 07810v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used as judges for evaluating model outputs, but their high cost, latency, and opacity limit scalability.
By Anish Laddha, Nitesh Pradhan, Gaurav Srivastava
arXiv:2609.24516v1 Announce Type: new
Abstract: In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these s...
By Khaoula Chehbouni, Melina Medjdoub, Florian Carichon, Golnoosh Farnadi, Jackie Chi Kit Cheung
arXiv:2606. 13685v1 Announce Type: cross Abstract: LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized.
By Abel Yagubyan