arXiv:2608. 02455v1 Announce Type: new Abstract: Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth.
By Zejun Xie, Xintong Li, Guang Wang, Desheng Zhang
Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features.
SCOPE is a framework that calibrates an acceptance threshold for large language models used as pairwise judges, ensuring that the error rate among non-abstained judgments does not exceed a user-specified level α. It introduces Bidirectional Preference Entropy (BPE) to provide a bias-neutral uncertainty signal by querying the judge in both response positions and converting the averaged preference probability into an entropy-based score. Across multiple pairwise judging benchmarks, BPE outperforms standard confidence proxies in calibration and discrimination, while SCOPE consistently meets the target risk bound (empirical FDR ≈0.097–0.099 at α=0.10) and retains substantial coverage, accepting up to 2.4× more judgments under the same risk constraint.
By Sher Badshah, Ali Emami, Hassan Sajjad
arXiv:2606. 13221v2 Announce Type: replace Abstract: Evaluating new large language models typically requires costly human annotation campaigns at scale.
By Bora Kargi, David Salinas
arXiv:2606. 13685v1 Announce Type: cross Abstract: LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized.
By Abel Yagubyan
arXiv:2607. 25136v1 Announce Type: new Abstract: Research on preference optimization often varies the training objective while holding the data fixed.
By Zhengtao Yao, Runhao Li, Xupeng Chen, Jiayi Cheng, Chenqian Le, Michael Yue, Siheng Wang, Haoyan Xu, Yuqi Li, Chenhao Wei, Zhengdao Li, Rongchao Zhang, Guang Yang, Yidong Wang, Junhao Dong