arXiv:2608. 02081v1 Announce Type: new Abstract: In this paper, we study prediction problems for paired comparison data, for example, predicting the win probability between two unmatched players and ranking all the players according to the order of their strengths by using win probability data between two matched players.
By Ryoya Yamasaki
The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.
By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv:2608. 10045v1 Announce Type: cross Abstract: The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models.
By Kaustubh Shivshankar Shejole, Tanish Agarwal, Arpit Agarwal, Avishek Ghosh
arXiv:2608.25200v2 Announce Type: replace-cross
Abstract: We study learning a mixture of $k$ Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous und...
By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang
arXiv:2604. 17805v2 Announce Type: replace-cross Abstract: Pairwise ranking systems based on Maximum Likelihood Estimation (MLE), such as the Bradley-Terry model, are widely used to aggregate preferences from pairwise comparisons.
By Junyi Yao, Zihao Zheng, Jiayu Long
The paper treats large language model (LLM) evaluation as a tensor completion problem, modeling noisy, sparse, and non‑uniform pairwise human judgments through a low‑rank latent score tensor under Bradley‑Terry‑Luce‑type models. It derives the efficient influence function and semiparametric efficiency bound for smooth functionals of the true tensor, and proposes a one‑step debiased estimator with asymptotic normality. A key innovation is a score‑whitening technique that equalizes local Fisher information, overcoming anisotropy in the information operator and enabling stable inference at optimal sample‑complexity.
By Jiachun Li, David Simchi-Levi, Will Wei Sun