arXiv Computation and Language

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

The paper investigates how well large language models (LLMs) can serve as judges in summarization evaluation by applying psychometric techniques. Using Many‑Facet Rasch Models, the authors decompose human and LLM ratings into latent summary quality, rater severity, dimension severity, and rating‑scale thresholds, and introduce a residual hardness metric to capture judging difficulty. Their analysis of 17 open‑weight LLM judges on the SummEval dataset reveals that moderate alignment in latent quality does not translate to alignment in residual hardness; humans and LLMs differ in which summary–dimension units are hard, with LLMs tending to find consistency hard and humans tending to find coherence hard, and some hard cases can be predicted from observable source–summary properties.

arXiv AI
Sep 3

Rating the Raters: Rasch Measurement Theory for LLM Evaluation

The paper discusses how large language models (LLMs) are used in various evaluation roles—examining benchmarks, judging other models, and rating human content—and frames each as a measurement problem. It proposes using Rasch measurement theory (RMT) to decompose ordinal ratings into distinct facets on a common scale, offering diagnostics for miscalibration and rater bias. A case study applying RMT to the Measuring Hate Speech corpus reveals systematic differences between LLMs and human raters in severity, calibration, robustness, sensitivity, and scale use, suggesting RMT should be part of the evaluation toolkit for LLMs in all roles.

By Pratik S. Sachdeva, Nathan Boudol
arXiv Computation and Language
Sep 3

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

The paper investigates whether consensus among large language model (LLM) judges truly reflects human alignment. By treating each judge’s scores as vectors, the authors measure spread, effective rank, and angles to human scores across 42 judges on Indic benchmarks, revealing that inter‑judge agreement often mirrors shared blind spots rather than human judgments. They find that while judges agree as much as humans, they only reach 58‑66% of human agreement and frequently focus on axes humans do not weight, indicating that ensemble agreement alone is insufficient evidence of alignment.

By Sourabrata Mukherjee, Hamna Hamna, Kalika Bali, Sunayana Sitaram
arXiv Machine Learning
Sep 21

How Many Humans Is a Judge Panel Worth?

arXiv:2609.21277v1 Announce Type: cross Abstract: How many human judgments does a panel of language models represent? The answer depends on what is matched. We audit categorical judge panels against...

By Chao Li, Yingying Yu, Yunfeng Li
Hugging Face Trending Papers
Aug 17

Toward Better Assessment of LLMs' Performance in Clinical Error Detection

Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart.