arXiv Machine Learning By Bitya Neuhof, Yuval Benjamini

Quantifying Ranking Uncertainty in LLM Benchmarks

Read the original on arXiv Machine Learning →

arXiv:2607. 16259v1 Announce Type: new Abstract: Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Statistics ML
Sep 4

Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons

The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.

By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv Machine Learning
Sep 14

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

The paper discusses how to quantify statistical uncertainty for aggregate performance metrics in machine learning benchmarks, focusing on methods such as bootstrapping, Bayesian hierarchical modeling, and visualizing task weightings with standard errors. It demonstrates that these techniques can uncover insights—for example, revealing that a model may dominate specific task types even if its overall performance is poor. The authors apply their approach to the Visual Task Adaptation Benchmark (VTAB) to illustrate its practical usefulness.

By Rachel Longjohn, Giri Gopalan, Emily Casleton