arXiv AI By Anissa Alloula, Federico Licini, Ava Batchkala, Seraphina Goldfarb-Tarrant

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators

Read the original on arXiv AI →

arXiv:2606. 07874v1 Announce Type: new Abstract: LLMs-as-judges are the only way to evaluate safety at scale.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.

By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong