arXiv Machine Learning

Pairwise Comparisons without Stochastic Transitivity: Model, Theory and Applications

arXiv:2501. 07437v3 Announce Type: replace-cross Abstract: Most statistical models for pairwise comparisons, including the Bradley-Terry (BT) and Thurstone models and many extensions, make a relatively strong assumption of stochastic transitivity.

arXiv Machine Learning
Aug 4

Isotonic Bradley-Terry Model for Paired Comparison Data

arXiv:2608. 02081v1 Announce Type: new Abstract: In this paper, we study prediction problems for paired comparison data, for example, predicting the win probability between two unmatched players and ranking all the players according to the order of their strengths by using win probability data between two matched players.

By Ryoya Yamasaki
arXiv Statistics ML
Sep 4

Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons

The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.

By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv AI
Sep 4

LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency

The paper treats large language model (LLM) evaluation as a tensor completion problem, modeling noisy, sparse, and non‑uniform pairwise human judgments through a low‑rank latent score tensor under Bradley‑Terry‑Luce‑type models. It derives the efficient influence function and semiparametric efficiency bound for smooth functionals of the true tensor, and proposes a one‑step debiased estimator with asymptotic normality. A key innovation is a score‑whitening technique that equalizes local Fisher information, overcoming anisotropy in the information operator and enabling stable inference at optimal sample‑complexity.

By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv AI
Jul 21

Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals

arXiv:2602. 03061v2 Announce Type: replace-cross Abstract: Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable rankings across platforms.

By Zihan Dong, Zhixian Zhang, Yang Zhou, Can Jin, Ruijia Wu, Linjun Zhang