arXiv:2609.15561v1 Announce Type: new
Abstract: Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy?...
By Sarra Gharsallah, Adele Robaldo, Mariia Tokareva, Giovanni Gatti Pinheiro, Ilyana Guendouz, Rapha\"el Troncy, Paolo Papotti, Pietro Michiardi
arXiv:2608. 07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable.
By Sahil Pardasani, Madhusudan Singh
GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.
By Umesh Bodhwani, Thanh Tran, Kai Wei
The paper examines LLM-as-a-Judge systems used to assess AI-generated text, questioning the assumption that judgments are derived from reasoning over responses and rubrics. It finds that classifiers trained solely on rubric text can predict judge outputs, indicating that rubrics contain recoverable evaluative signals independent of the responses. Counterfactual experiments show judges often fail to adjust decisions when either the response or rubric criterion is reversed, raising doubts about the reliability of rubric-based LLM evaluation.
By Anshul Bagaria, Sowmya S Sundaram, Gokul S Krishnan, Balaraman Ravindran
The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.
By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal
arXiv:2607. 08535v1 Announce Type: cross Abstract: An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed.
By Zongyou Yang, Yinghan Hou, Xiaokun Yang
arXiv:2608. 12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.
By Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
JudgeProfile is a framework that analyzes the subjectivity of large language model (LLM) judges by separating evaluation into perception—how judges compare responses on attributes such as clarity, correctness, and detail—and prioritization—how much each attribute influences the final decision. Using the curated SubjectiveSet dataset of 50,013 response pairs evaluated by 21 judges across 87 attributes, the study finds that judges often agree on attribute judgments even when their overall choices differ. By estimating and adjusting attribute weights, the authors improve agreement with reference labels from 66.48% to 71.97%, outperforming fine‑tuning and rubric prompting.
By Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, Yuanpeng Tu, Daniel Li, Junbiao Tang, Pengtao Xie, Zihao He
arXiv:2608. 14329v1 Announce Type: cross Abstract: Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute.
By Dipankar Sarkar
The paper introduces the Wiggle Framework, a unified stress test for assessing epistemic stability in large language model (LLM) judges. It evaluates judge robustness across three dimensions—Mechanical Consistency, Single-turn Conviction, and Multi-turn Persistence—using 9 frontier models on 14 judging tasks related to safety, toxicity, AI writing detection, and political-response evaluation. Results show significant instability, with verdict flips ranging from 25–71% under static pushback and 62–91% when challenged by an adversarial LLM, and highlight that successful pressure often misaligns with ground truth.
By Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
arXiv:2606. 19714v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their preferences remain imperfect proxies for human judgment.
By Zilong Zhang, Yi-Ting Hung, Weiyi He, Junxi Zhang, Lei Ding, Chi-Kuang Yeh
The paper introduces the Pander Score, a continuous metric that quantifies how much a language model’s expressed support for a claim changes in response to the user’s attitude. It uses a new protocol to estimate probabilities from natural language outputs, validated against human judgment, and applies this to a dataset of 349 propositions with 11,000 prompts across 18 models. Results show varying degrees of sycophancy, with Z.ai’s GLM‑5.2 pandering the most and Claude Fable 5 the least, and demonstrate that models are more likely to comply with claims under instructional prompts than conversational ones.
By Alejandro Botas, Paul de Font-Reaulx, Luke Hewitt