arXiv Computation and Language

Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency

Large Language Model judges are commonly used to rank texts via pairwise comparison, with reliability traditionally measured by position bias, transitivity, and pairwise agreement. This paper argues that these proxies are misleading because they are dominated by close‑rank‑gap pairs, which contribute little to the overall ranking, while far‑gap pairs carry the true ranking signal. Experiments on simulations and human‑rated corpora show weak correlation between the proxies and actual ranking accuracy, suggesting judges should be evaluated using rank‑gap‑conditional metrics against human rankings.

arXiv AI
Jun 3

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

arXiv:2606. 03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks cannot be trusted, their items having likely leaked into pretraining, so scores reflect memorization rather than fitness.

By Alexander Apartsin, Yehudit Aperstein
Hugging Face Trending Papers
Aug 3

Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features.

arXiv Computation and Language
Sep 3

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

The paper investigates whether consensus among large language model (LLM) judges truly reflects human alignment. By treating each judge’s scores as vectors, the authors measure spread, effective rank, and angles to human scores across 42 judges on Indic benchmarks, revealing that inter‑judge agreement often mirrors shared blind spots rather than human judgments. They find that while judges agree as much as humans, they only reach 58‑66% of human agreement and frequently focus on axes humans do not weight, indicating that ensemble agreement alone is insufficient evidence of alignment.

By Sourabrata Mukherjee, Hamna Hamna, Kalika Bali, Sunayana Sitaram