arXiv Machine Learning

CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion

arXiv:2607. 05046v1 Announce Type: new Abstract: Evaluating generative AI models is a routine, but resource-intensive, process that is conducted over and over again during the course of model development.

arXiv AI
Sep 18

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

The paper introduces prediction‑powered smoothing (PP‑S) and its taxonomy‑aware extension (PP‑TS) to improve point and interval estimates of domain‑specific AI performance when only a limited sample of labeled units is available. It also proposes a new design‑based cross‑validation score that is approximately unbiased for selecting between direct and smoothed estimators. Experiments on a curated benchmark and real‑world agent traffic show that the proposed methods outperform direct estimators in both accuracy and coverage, and that the new score matches the performance of an independent validation sample while providing more precise error estimates.

By Sho Kawano, Zehang Richard Li, Paul A. Parker
arXiv AI
Jun 4

Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-Learning

arXiv:2605. 23595v2 Announce Type: replace-cross Abstract: The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data.

By Trinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen, Thanh Tam Nguyen
arXiv Machine Learning
Jun 24

You Don't Need to Run Every Eval

arXiv:2606. 24020v1 Announce Type: new Abstract: A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare design choices, and select the checkpoint for the release.

By Yuchen Zeng, Dimitris Papailiopoulos
arXiv Computation and Language
Aug 28

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

The paper introduces prediction‑powered evaluation, a framework that blends limited human judgments with large‑scale automatic scores to produce unbiased, data‑efficient system comparisons. It offers both parametric and non‑parametric methods, examines the trade‑off between paired and unpaired designs, and validates the approach on six WMT datasets. Additionally, the authors propose the Prediction‑Powered Saving Ratio (PPSR), a meta‑metric that quantifies how much human annotation can be saved by using an automatic metric within this framework, providing more discriminative and stable metric rankings than existing system‑level meta‑metrics.

By Mingqi Gao, Anthony Sicilia, Weiyan Shi
arXiv AI
Jul 3

PACE: A Proxy for Agentic Capability Evaluation

arXiv:2607. 02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure.

By Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, Graham Neubig