arXiv Machine Learning By Abhishek Divekar

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

Read the original on arXiv Machine Learning →

arXiv:2606. 05308v1 Announce Type: new Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 28

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

The paper introduces prediction‑powered evaluation, a framework that blends limited human judgments with large‑scale automatic scores to produce unbiased, data‑efficient system comparisons. It offers both parametric and non‑parametric methods, examines the trade‑off between paired and unpaired designs, and validates the approach on six WMT datasets. Additionally, the authors propose the Prediction‑Powered Saving Ratio (PPSR), a meta‑metric that quantifies how much human annotation can be saved by using an automatic metric within this framework, providing more discriminative and stable metric rankings than existing system‑level meta‑metrics.

By Mingqi Gao, Anthony Sicilia, Weiyan Shi