arXiv:2501. 07437v3 Announce Type: replace-cross Abstract: Most statistical models for pairwise comparisons, including the Bradley-Terry (BT) and Thurstone models and many extensions, make a relatively strong assumption of stochastic transitivity.
By Sze Ming Lee, Yunxiao Chen
arXiv:2510. 20454v2 Announce Type: replace Abstract: Intransitive player dominance, where player A beats B, B beats C, but C beats A, is common in competitive tennis.
By Lawrence Clegg, John Cartlidge
arXiv:2505. 23437v2 Announce Type: replace-cross Abstract: Ranking systems influence decision-making in high-stakes domains like health, education, and employment, where they can have substantial economic and social impacts.
By Antonio Ferrara, Andrea Pugnana, Francesco Bonchi, Salvatore Ruggieri
arXiv:2604. 17805v2 Announce Type: replace-cross Abstract: Pairwise ranking systems based on Maximum Likelihood Estimation (MLE), such as the Bradley-Terry model, are widely used to aggregate preferences from pairwise comparisons.
By Junyi Yao, Zihao Zheng, Jiayu Long
arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.
By Dennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan Feuerriegel
arXiv:2506. 06989v3 Announce Type: replace-cross Abstract: Learning-to-rank (LTR) systems commonly depend on implicit feedback, such as user clicks, because it is easy to collect and can serve as a valuable signal of user preferences.
By Md Aminul Islam, Kathryn Vasilaky, Elena Zheleva
arXiv:2609.07617v1 Announce Type: new
Abstract: With the rise of live sports betting in recent years, tennis forecasting has expanded from pre-match prediction to models that update win probabilities...
By Charles Xie, Aneesh Muppidi
arXiv:2608.25200v2 Announce Type: replace-cross
Abstract: We study learning a mixture of $k$ Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous und...
By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang
arXiv:2608. 03437v1 Announce Type: cross Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability.
By Vil\'em Zouhar, Julia Kreutzer, Alon Lavie, Tom Kocmi, Matt Post, Ond\v{r}ej Bojar, Mrinmaya Sachan
arXiv:2606. 06043v1 Announce Type: cross Abstract: Follow-the-regularized-leader framework has shown effectiveness and flexibility in online learning problems, where the choice of learning rates are known to be crucial.
By Jongyeong Lee, Junya Honda, Shinji Ito, Chansoo Kim
The paper introduces MoPLEx, an expectation‑maximization algorithm for learning mixtures of Plackett‑Luce models from multi‑way ranking data. It augments rankings with synthetic responses from a base language model and uses a gradient‑based estimation to reduce inference cost, enabling efficient fitting of large‑scale models. Experiments show the method achieves low probability estimation error and improves clustering and ranking accuracy by 43.7% and 15.2% over baselines.
By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang
arXiv:2609.36740v1 Announce Type: new
Abstract: Many recommender systems such as for e-commerce and news platforms aim to provide users with rankings they are likely to interact with. Off-Policy Lear...
By Ren Kishimoto, Koichi Tanaka, Haruka Kiyohara, Yusuke Narita, Yasuo Yamamoto, Nobuyuki Shimizu, Yuta Saito