Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation
Read the original on arXiv Computation and Language →The paper introduces prediction‑powered evaluation, a framework that blends limited human judgments with large‑scale automatic scores to produce unbiased, data‑efficient system comparisons. It offers both parametric and non‑parametric methods, examines the trade‑off between paired and unpaired designs, and validates the approach on six WMT datasets. Additionally, the authors propose the Prediction‑Powered Saving Ratio (PPSR), a meta‑metric that quantifies how much human annotation can be saved by using an automatic metric within this framework, providing more discriminative and stable metric rankings than existing system‑level meta‑metrics.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.