arXiv Computation and Language By Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani

Post-hoc Alignment of LLM-judges to Human Judgment Distribution

Read the original on arXiv Computation and Language →

The paper introduces NAPHA, a lightweight post‑hoc alignment method that improves large language model (LLM) predictions of human judgment distributions (HJD) by matching LLM output distributions to HJD through entropy‑based class assignment and specialized alignment models. Experiments on five datasets show that while LLMs perform near human‑level on hard‑label tasks, they struggle with soft‑label predictions, and NAPHA consistently enhances soft‑label accuracy, especially on high‑entropy instances. The study also demonstrates that better entropy class prediction can further boost NAPHA’s effectiveness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 16

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

The paper introduces JudgeBiasBench, a benchmark that systematically quantifies judgment biases in large language model (LLM)-based judges across four dimensions and 12 bias types. It evaluates both generative and discriminative judges, revealing significant bias patterns that undermine reliability. The authors propose bias-aware training—reinforcement learning for generative judges and contrastive learning for discriminative judges—to reduce these biases while maintaining evaluation performance.

By Hongli Zhou, Hui Huang, Rui Zhang, Kehai Chen, Bing Xu, Conghui Zhu, Tiejun Zhao, Muyun Yang
arXiv Machine Learning
Sep 14

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.

By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal