arXiv AI

From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks

The paper investigates how the design of LLM-as-a-Judge protocols influences both the intrinsic quality of judgments and their downstream utility in open-ended tasks. By varying verdict granularity, critique usage, and evaluation batching, and by applying Judge guidance to test-time inference methods such as Best-of-N selection, revision, and beam search, the authors find that judgment quality and downstream performance do not always align and that protocol choices significantly affect outcomes. The study highlights the need for comprehensive evaluation of LLM Judges that considers both judgment quality and practical utility.

arXiv AI
Aug 20

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

The paper introduces SESSE, a training‑free framework that breaks down LLM‑as‑judge evaluations into five steps—Sketch, Expand, Sort, Summarize, Evaluate—by mining sub‑questions from the judge’s own error cases. It requires no oracle responses, task‑specific rubrics, or fine‑tuning, yet on RewardBench it matches the performance of chain‑of‑thought baselines and rivals a fine‑tuned specialist (RISE‑Judge‑32B). SESSE provides per‑criterion vote evidence, offering an interpretable audit trail that can diagnose label ambiguity and judge failure modes that a single holistic output token cannot reveal.

By Dae Lee, Mihai Delgeanu, Adel Youssef
Hugging Face Trending Papers
Jun 3

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is largely present before any targeted training: prompted few-shot, a base model already predicts an external judge's multi-attribute quality scores on open-ended responses well above chance across three benchmarks.

arXiv Computation and Language
Sep 16

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

The paper introduces JudgeBiasBench, a benchmark that systematically quantifies judgment biases in large language model (LLM)-based judges across four dimensions and 12 bias types. It evaluates both generative and discriminative judges, revealing significant bias patterns that undermine reliability. The authors propose bias-aware training—reinforcement learning for generative judges and contrastive learning for discriminative judges—to reduce these biases while maintaining evaluation performance.

By Hongli Zhou, Hui Huang, Rui Zhang, Kehai Chen, Bing Xu, Conghui Zhu, Tiejun Zhao, Muyun Yang
arXiv AI
3d ago

JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

JudgeProfile is a framework that analyzes the subjectivity of large language model (LLM) judges by separating evaluation into perception—how judges compare responses on attributes such as clarity, correctness, and detail—and prioritization—how much each attribute influences the final decision. Using the curated SubjectiveSet dataset of 50,013 response pairs evaluated by 21 judges across 87 attributes, the study finds that judges often agree on attribute judgments even when their overall choices differ. By estimating and adjusting attribute weights, the authors improve agreement with reference labels from 66.48% to 71.97%, outperforming fine‑tuning and rubric prompting.

By Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, Yuanpeng Tu, Daniel Li, Junbiao Tang, Pengtao Xie, Zihao He
arXiv Machine Learning
Sep 14

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.

By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal
arXiv AI
Jun 3

JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

arXiv:2605. 25240v2 Announce Type: replace-cross Abstract: Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs.

By Russell Yang, Ruishi Chen, Pierce Kelaita, Riya Ranjan, Sibo Ma, Charles Dickens, Matthew Guillod, Megan Ma, Julian Nyarko
Hugging Face Trending Papers
Aug 18

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

The paper introduces a risk‑controlled framework for using large language models (LLMs) as judges in tasks without reference answers. By calibrating uncertainty thresholds on a held‑out set, the method ensures that the false discovery rate of accepted verdicts stays below a user‑specified level α with high probability, using finite‑sample Clopper–Pearson intervals. When the parametric judge lacks confidence, the instance is routed to a retrieval‑augmented mode with a second calibrated threshold, preserving the error guarantee while achieving higher coverage than single‑mode baselines.