arXiv AI

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

arXiv:2606. 09409v1 Announce Type: new Abstract: Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge biases.

arXiv Computation and Language
4d ago

Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency

Large Language Model judges are commonly used to rank texts via pairwise comparison, with reliability traditionally measured by position bias, transitivity, and pairwise agreement. This paper argues that these proxies are misleading because they are dominated by close‑rank‑gap pairs, which contribute little to the overall ranking, while far‑gap pairs carry the true ranking signal. Experiments on simulations and human‑rated corpora show weak correlation between the proxies and actual ranking accuracy, suggesting judges should be evaluated using rank‑gap‑conditional metrics against human rankings.

By Bruno Brocai, Maria Becker
arXiv AI
Aug 20

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

The study investigates bias in large language model (LLM) judges by having ten LLMs evaluate narrative constraint selections rather than generated text. Results show that self-preference largely disappears under blind evaluation when quality and evaluator severity are controlled, but self- and other-labels alone shift scores bidirectionally when quality is matched. The authors conclude that authorship attribution drives evaluation bias and that open-ended, ground‑truth‑free tasks can effectively study LLM judge behavior.

By Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
arXiv AI
6d ago

Accounting for Bias Enables Sustainable LLM Evaluation

The paper argues that the current LLM-as-a-judge evaluation method, which compensates for systematic measurement bias by increasing the number of comparisons, is statistically unsound and computationally wasteful. It identifies that treating LLM judges as neutral ignores documented biases such as position bias, verbosity bias, judge severity, and self‑enhancement. The authors propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, enabling reliable rankings with far fewer comparisons and negligible additional compute.

By Harshita Katoch, David Antony Selby, Gerrit Gro{\ss}mann, Sebastian Vollmer
arXiv Computation and Language
Sep 15

Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation

arXiv:2609.15561v1 Announce Type: new Abstract: Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy?...

By Sarra Gharsallah, Adele Robaldo, Mariia Tokareva, Giovanni Gatti Pinheiro, Ilyana Guendouz, Rapha\"el Troncy, Paolo Papotti, Pietro Michiardi
arXiv AI
4d ago

JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

JudgeProfile is a framework that analyzes the subjectivity of large language model (LLM) judges by separating evaluation into perception—how judges compare responses on attributes such as clarity, correctness, and detail—and prioritization—how much each attribute influences the final decision. Using the curated SubjectiveSet dataset of 50,013 response pairs evaluated by 21 judges across 87 attributes, the study finds that judges often agree on attribute judgments even when their overall choices differ. By estimating and adjusting attribute weights, the authors improve agreement with reference labels from 66.48% to 71.97%, outperforming fine‑tuning and rubric prompting.

By Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, Yuanpeng Tu, Daniel Li, Junbiao Tang, Pengtao Xie, Zihao He
arXiv AI
Sep 4

Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

The paper examines LLM-as-a-Judge systems used to assess AI-generated text, questioning the assumption that judgments are derived from reasoning over responses and rubrics. It finds that classifiers trained solely on rubric text can predict judge outputs, indicating that rubrics contain recoverable evaluative signals independent of the responses. Counterfactual experiments show judges often fail to adjust decisions when either the response or rubric criterion is reversed, raising doubts about the reliability of rubric-based LLM evaluation.

By Anshul Bagaria, Sowmya S Sundaram, Gokul S Krishnan, Balaraman Ravindran