arXiv:2607. 08535v1 Announce Type: cross Abstract: An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed.
By Zongyou Yang, Yinghan Hou, Xiaokun Yang
The paper demonstrates that reading a large language model (LLM) judge’s verdict from the logits of its first generated token—an approach used in constrained decoding and likelihood‑scoring evaluation—introduces a significant distortion in position bias. Because judges do not always start with a verdict token (12–49% of cases for Qwen3 judges and <3% for Llama‑3.1‑8B and Phi‑3.5‑mini), this readout often returns the first response rather than a true judgment, inflating position bias by up to 42 points while barely affecting judge accuracy. The authors recommend reporting the frequency with which a judge leads with a verdict token to provide a more accurate assessment of position bias.
By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.
By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal
The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.
By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
arXiv:2608. 12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.
By Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
The paper introduces the Wiggle Framework, a unified stress test for assessing epistemic stability in large language model (LLM) judges. It evaluates judge robustness across three dimensions—Mechanical Consistency, Single-turn Conviction, and Multi-turn Persistence—using 9 frontier models on 14 judging tasks related to safety, toxicity, AI writing detection, and political-response evaluation. Results show significant instability, with verdict flips ranging from 25–71% under static pushback and 62–91% when challenged by an adversarial LLM, and highlight that successful pressure often misaligns with ground truth.
By Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
The paper introduces TrustSwap, a counterfactual test that swaps or removes source reliability labels while keeping evidence text constant, to evaluate how retrieval‑augmented fact‑checking models respond across verdict, confidence, and search decisions. Experiments on untrained and RL‑trained models show that confidence and search largely follow labels, yet label changes can flip a significant portion of verdicts, especially in larger models. The authors propose trust‑swap augmentation (TSA) to mitigate this shortcut, demonstrating reduced verdict flip rates and maintained accuracy in several settings, though its effectiveness diminishes at larger model scales.
By Jianchang Su, Yiwei Yang, Wei Zhang
arXiv:2610.07177v1 Announce Type: cross
Abstract: An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of...
By Gowthamkumar Nandakishore
arXiv:2606. 13685v1 Announce Type: cross Abstract: LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized.
By Abel Yagubyan
arXiv:2606. 05384v1 Announce Type: new Abstract: LLM-as-judge evaluation is widely used in benchmarking pipelines, where model outputs are compared and ranked using automated evaluators.
By Srimonti Dutta, Akshata Kishore Moharir
arXiv:2606. 25487v1 Announce Type: cross Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade.
By Yang Gao (Veyon Solutions)
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade. The judge is rarely checked.