arXiv Computation and Language

LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora

arXiv AI
Aug 26

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.

By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
arXiv AI
1d ago

Rating the Raters: Rasch Measurement Theory for LLM Evaluation

The paper discusses how large language models (LLMs) are used in various evaluation roles—examining benchmarks, judging other models, and rating human content—and frames each as a measurement problem. It proposes using Rasch measurement theory (RMT) to decompose ordinal ratings into distinct facets on a common scale, offering diagnostics for miscalibration and rater bias. A case study applying RMT to the Measuring Hate Speech corpus reveals systematic differences between LLMs and human raters in severity, calibration, robustness, sensitivity, and scale use, suggesting RMT should be part of the evaluation toolkit for LLMs in all roles.

By Pratik S. Sachdeva, Nathan Boudol
arXiv Machine Learning
Aug 28

Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores

The paper presents an interpretable, fair, and accurately benchmarked automated system for assessing second‑language English speaking. Using a hybrid of feature‑based speech‑timing metrics and a large language model (LLM) fluency judgment, the system achieves a Spearman correlation of 0.818 with the ICNALE Global Rating Archive, outperforming 81 % of trained human raters. A controlled study shows that encoding pauses into the LLM prompt does not meaningfully affect fluency scores, indicating that the system’s fluency signal derives from measurable speech‑timing features.

By Eichi Uehara