arXiv Machine Learning

Three Ways Classical Test Theory Misleads for LLM Judges

The article examines how classical test theory reliability statistics misrepresent the performance of large language model (LLM) judges. It shows that internal‑consistency coefficients, the dependability index, and Livingston‑Lewis accuracy each conflate judge error with item design or criterion validity, making it impossible to attribute a single reliability value to the judge alone. The authors argue that such misattribution can influence deployment decisions and documentation.

arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand
arXiv AI
Aug 26

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.

By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
arXiv AI
Sep 2

Commit-first LLM judging inherits the judge's own errors

The paper investigates whether widely used evaluation frameworks for large language models (LLMs) implement a defense called commit‑first judging, which requires a judge to solve a task itself before accepting a candidate answer. Across 24 configurations in eight popular frameworks, none use the full commit‑first method; nine use a weaker variant that is ineffective. In controlled experiments, the weaker variant allowed systems to game the judge, while the full commit‑first approach eliminated this vulnerability but sometimes worsened evaluation when the judge’s own answer was incorrect.

By Idil Gozel
arXiv Computation and Language
6d ago

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

The study investigates whether Jev, a typed classifier that outputs probabilities over allowed answers without generating text, can replace large language model (LLM) rubric judges. Across nine panels from seven benchmarks, Jev’s accuracy differed significantly from LLM judges in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons were inconclusive. In terms of cost and speed, Jev was 29 to 325 times cheaper and 30 to 220 times faster than the flash‑tier LLM judges, and a cascade approach that defers uncertain Jev verdicts to an LLM yielded only modest gains. whyItMatters":"The findings suggest that a lightweight classifier like Jev can serve as an efficient first‑stage evaluator, potentially reducing the reliance on expensive and slow LLM judges in automated grading pipelines."

By Delip Rao, Chris Callison-Burch
arXiv Machine Learning
Sep 10

A Closed-Form Estimator and Diagnostic Battery for Anchor-Judge Error Correlation, Under a Single-Common-Factor Model

The paper presents a closed‑form estimator for the contamination correlation between anchors and judges under a single‑common‑factor model, requiring at least two judges and two anchors. It introduces a diagnostic battery—including judge‑covariance dispersion, over‑identification tests, a family‑block test, bootstrap confidence intervals, and a weak‑identification screen—to validate the estimator and detect violations. The authors also discuss identification limits for ordinal data and report that real panels have not yet passed the model‑adequacy pre‑test, while simulation studies confirm the estimator’s performance.

By Veerendra Kumar Sunkavalli
Hugging Face Trending Papers
6d ago

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

The study compares Jev, a typed classifier that outputs probabilities over allowed answers, with three flash-tier LLM rubric judges across nine panels from seven benchmarks. Jev’s accuracy differs significantly from an LLM judge in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons are inconclusive. In terms of cost and speed, Jev is 29 to 325 times cheaper and 30 to 220 times faster than the LLM judges, and a cascade that defers uncertain Jev verdicts to an LLM yields only modest gains of up to 2.0 points over the best single judge.