arXiv Machine Learning By Bora Kargi, David Salinas

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation

Read the original on arXiv Machine Learning →

arXiv:2606. 13221v2 Announce Type: replace Abstract: Evaluating new large language models typically requires costly human annotation campaigns at scale.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 3

Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features.