arXiv Computation and Language By Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

Read the original on arXiv Computation and Language →

The paper investigates how well large language models (LLMs) can serve as judges in summarization evaluation by applying psychometric techniques. Using Many‑Facet Rasch Models, the authors decompose human and LLM ratings into latent summary quality, rater severity, dimension severity, and rating‑scale thresholds, and introduce a residual hardness metric to capture judging difficulty. Their analysis of 17 open‑weight LLM judges on the SummEval dataset reveals that moderate alignment in latent quality does not translate to alignment in residual hardness; humans and LLMs differ in which summary–dimension units are hard, with LLMs tending to find consistency hard and humans tending to find coherence hard, and some hard cases can be predicted from observable source–summary properties.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 3

Rating the Raters: Rasch Measurement Theory for LLM Evaluation

The paper discusses how large language models (LLMs) are used in various evaluation roles—examining benchmarks, judging other models, and rating human content—and frames each as a measurement problem. It proposes using Rasch measurement theory (RMT) to decompose ordinal ratings into distinct facets on a common scale, offering diagnostics for miscalibration and rater bias. A case study applying RMT to the Measuring Hate Speech corpus reveals systematic differences between LLMs and human raters in severity, calibration, robustness, sensitivity, and scale use, suggesting RMT should be part of the evaluation toolkit for LLMs in all roles.

By Pratik S. Sachdeva, Nathan Boudol