Decomposing LLM-Judge Uncertainty to Target Expert Labels
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In p...
arXiv:2607. 20529v1 Announce Type: cross Abstract: Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs.
The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.
The paper introduces a method for learning from multiple experts who provide interval labels, addressing both within‑label imprecision and between‑expert variation. It harmonizes diverse label vocabularies into a shared probabilistic space, retains individual intervals using a mixture of Beta distributions, and decomposes predictive uncertainty into components that are matched to their corresponding sources of label uncertainty. On sea‑ice concentration data, the approach achieves a 31% reduction in mean absolute error compared to hard‑label baselines and outperforms several aggregation and interval‑regression methods.
arXiv:2609.26550v3 Announce Type: replace Abstract: LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge tha...
JudgeProfile is a framework that analyzes the subjectivity of large language model (LLM) judges by separating evaluation into perception—how judges compare responses on attributes such as clarity, correctness, and detail—and prioritization—how much each attribute influences the final decision. Using the curated SubjectiveSet dataset of 50,013 response pairs evaluated by 21 judges across 87 attributes, the study finds that judges often agree on attribute judgments even when their overall choices differ. By estimating and adjusting attribute weights, the authors improve agreement with reference labels from 66.48% to 71.97%, outperforming fine‑tuning and rubric prompting.