arXiv:2607. 05046v1 Announce Type: new Abstract: Evaluating generative AI models is a routine, but resource-intensive, process that is conducted over and over again during the course of model development.
By Adam Fisch, Daniel Deutsch, Joshua Maynez, Alekh Agarwal, Jonathan Berant, William Cohen, Amir Globerson, Jacob Eisenstein
arXiv:2604. 23099v2 Announce Type: replace-cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks.
By Yizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi Wang
arXiv:2607. 16239v1 Announce Type: new Abstract: AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains.
By Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha
The paper introduces prediction‑powered evaluation, a framework that blends limited human judgments with large‑scale automatic scores to produce unbiased, data‑efficient system comparisons. It offers both parametric and non‑parametric methods, examines the trade‑off between paired and unpaired designs, and validates the approach on six WMT datasets. Additionally, the authors propose the Prediction‑Powered Saving Ratio (PPSR), a meta‑metric that quantifies how much human annotation can be saved by using an automatic metric within this framework, providing more discriminative and stable metric rankings than existing system‑level meta‑metrics.
By Mingqi Gao, Anthony Sicilia, Weiyan Shi
arXiv:2609.35815v1 Announce Type: cross
Abstract: Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated co...
By Ian Arawjo
arXiv:2606. 29784v1 Announce Type: cross Abstract: Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity.
By Xinrui Ruan, Zhenyu Zhao, Waverly Wei, Yueshan Zhang, Zeyu Zheng, Sui Huang, Jingshen Wang