arXiv Machine Learning By Abhishek Divekar

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

Read the original on arXiv Machine Learning →

arXiv:2606. 05308v1 Announce Type: new Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.