arXiv Machine Learning By Sze Ming Lee, Yunxiao Chen

Pairwise Comparisons without Stochastic Transitivity: Model, Theory and Applications

Read the original on arXiv Machine Learning →

arXiv:2501. 07437v3 Announce Type: replace-cross Abstract: Most statistical models for pairwise comparisons, including the Bradley-Terry (BT) and Thurstone models and many extensions, make a relatively strong assumption of stochastic transitivity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 4

Isotonic Bradley-Terry Model for Paired Comparison Data

arXiv:2608. 02081v1 Announce Type: new Abstract: In this paper, we study prediction problems for paired comparison data, for example, predicting the win probability between two unmatched players and ranking all the players according to the order of their strengths by using win probability data between two matched players.

By Ryoya Yamasaki
arXiv Statistics ML
Sep 4

Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons

The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.

By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv AI
Sep 4

LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency

The paper treats large language model (LLM) evaluation as a tensor completion problem, modeling noisy, sparse, and non‑uniform pairwise human judgments through a low‑rank latent score tensor under Bradley‑Terry‑Luce‑type models. It derives the efficient influence function and semiparametric efficiency bound for smooth functionals of the true tensor, and proposes a one‑step debiased estimator with asymptotic normality. A key innovation is a score‑whitening technique that equalizes local Fisher information, overcoming anisotropy in the information operator and enabling stable inference at optimal sample‑complexity.

By Jiachun Li, David Simchi-Levi, Will Wei Sun