Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation
arXiv:2608. 14684v1 Announce Type: cross Abstract: LLM judges increasingly evaluate responses against fine-grained rubric checklists.
arXiv:2606. 09165v1 Announce Type: new Abstract: Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they remain brittle under prompt and rubric variation, with false negative-rate swings of up to 0.
arXiv:2608. 14684v1 Announce Type: cross Abstract: LLM judges increasingly evaluate responses against fine-grained rubric checklists.
The paper examines LLM-as-a-Judge systems used to assess AI-generated text, questioning the assumption that judgments are derived from reasoning over responses and rubrics. It finds that classifiers trained solely on rubric text can predict judge outputs, indicating that rubrics contain recoverable evaluative signals independent of the responses. Counterfactual experiments show judges often fail to adjust decisions when either the response or rubric criterion is reversed, raising doubts about the reliability of rubric-based LLM evaluation.
arXiv:2601.08654v3 Announce Type: replace Abstract: Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the...
arXiv:2602.13576v2 Announce Type: replace-cross Abstract: Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-langu...
arXiv:2607. 01153v3 Announce Type: replace-cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task.
arXiv:2606. 09118v1 Announce Type: new Abstract: As LLM capabilities advance rapidly, the evaluation methods used to assess them increasingly lag behind.
arXiv:2603. 00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks.
The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.
arXiv:2602.02219v3 Announce Type: replace Abstract: Large language models are widely employed as evaluators, a paradigm commonly referred to as LLM-as-a-judge. Prior research has predominantly examin...
The paper introduces Evidence‑Diagnosed Intervention Training (EDIT), a two‑phase framework designed to improve rubric‑faithful grading by large language models. EDIT‑SFT first identifies problematic reasoning steps using internal model signals—posterior belief over the final mark and input‑grounding scores—and revises only those steps with rubric checklists. EDIT‑RL then calibrates the grader with belief‑guided reward shaping, penalising harmful belief drifts while encouraging useful exploration. Experiments on two real‑world, multi‑subject grading benchmarks show that EDIT consistently outperforms strong supervised fine‑tuning and reinforcement learning baselines, with ablation studies confirming the importance of internal‑state diagnostics.
The paper investigates whether automatic safety judges evaluate the content of a model’s reply or merely its style. By keeping the reply content fixed and adding various style wrappers—such as educational disclaimers, fake reasoning blocks, or token refusals—the authors show that many judges flip their verdicts, indicating that style can influence safety judgments. The study evaluates over 600 jailbreak examples across multiple judges, revealing that some judges are highly susceptible to style-based manipulation while others remain robust.
arXiv:2609.13824v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faste...