arXiv Machine Learning By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao

When Can You Debias an LLM Judge? Identifiability Limits, a Test, and Designs for Top-k Ranking

Read the original on arXiv Machine Learning →

arXiv:2607. 02104v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand
arXiv AI
1d ago

The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

The paper demonstrates that reading a large language model (LLM) judge’s verdict from the logits of its first generated token—an approach used in constrained decoding and likelihood‑scoring evaluation—introduces a significant distortion in position bias. Because judges do not always start with a verdict token (12–49% of cases for Qwen3 judges and <3% for Llama‑3.1‑8B and Phi‑3.5‑mini), this readout often returns the first response rather than a true judgment, inflating position bias by up to 42 points while barely affecting judge accuracy. The authors recommend reporting the frequency with which a judge leads with a verdict token to provide a more accurate assessment of position bias.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli