Large language models (LLMs) are increasingly used to evaluate output quality, but guaranteeing agreement with human judgments is difficult. The paper introduces a Localize-Then-Decide framework that first uses conformal prediction to narrow down a shortlist likely to contain the human-preferred response, then applies a calibrated confidence rule to select a single response or abstain. Experiments show this two-stage approach consistently yields higher guarantee success rates and greater coverage than single-stage baselines across various candidate sizes and datasets.
By Xinyu Li, Yi Zhou, Guanqun Cao, Zeyu Fu, Tianjin Huang, Gaojie Jin
arXiv:2609.38860v1 Announce Type: cross
Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference lear...
By Zhongman Du, Huiming Zhang, Haodong Zhu, Baochang Zhang
The paper introduces multi-expert Conformal Risk Control (CRC) algorithms for pairwise LLM-as-a-Judge evaluation in open-ended dialogue. Two initial methods—Score Averaging and Decision Voting—aggregate at the score and decision levels, respectively, and outperform single-expert approaches on homogeneous expert panels. To address limited coverage on heterogeneous panels, the authors propose Marginal‑Calibrated Conformal Consensus (MC3), which captures distinct per‑expert scoring scales through threshold ratios while maintaining a unified decision function, and demonstrate its effectiveness on a new 1,800‑pair human pairwise‑preference benchmark called Panel.
By Ming Cheng, Yusheng Dai, Qiuhong Ke, Zhaolin Chen, Lizhen Qu
arXiv:2606. 13221v2 Announce Type: replace Abstract: Evaluating new large language models typically requires costly human annotation campaigns at scale.
By Bora Kargi, David Salinas
arXiv:2602.02219v3 Announce Type: replace
Abstract: Large language models are widely employed as evaluators, a paradigm commonly referred to as LLM-as-a-judge. Prior research has predominantly examin...
By Yuzheng Xu, Tosho Hirasawa, Tadashi Kozuno, Yoshitaka Ushiku
The paper demonstrates that large language model (LLM) evaluators, whether reward‑model based or prompted LLM‑as‑a‑Judge, exhibit significant language bias in multilingual settings. Experiments with semantically identical instruction‑response pairs across 23 languages reveal that lower‑resource languages receive higher scores, a bias that persists across eight open‑weight evaluators and is not detectable by standard pairwise accuracy metrics. The authors link the bias to model uncertainty and language identity, showing it cannot be explained by content difficulty alone.
By Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen