arXiv Machine Learning

Nonparametric LLM Evaluation from Preference Data

arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.

arXiv Machine Learning
Aug 27

Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment

The paper introduces MoPLEx, an expectation‑maximization algorithm for learning mixtures of Plackett‑Luce models from multi‑way ranking data. It augments rankings with synthetic responses from a base language model and uses a gradient‑based estimation to reduce inference cost, enabling efficient fitting of large‑scale models. Experiments show the method achieves low probability estimation error and improves clustering and ranking accuracy by 43.7% and 15.2% over baselines.

By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang
arXiv AI
6d ago

DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration

The paper introduces DIAL, a framework that uses large language models (LLMs) as judges while mitigating position bias and aligning their judgments with human preferences. DIAL separates judge‑specific position effects, learns shared structure in debiased LLM preferences, and adaptively calibrates this structure toward human targets using limited human comparisons. Experiments on simulations and three human‑preference benchmarks show that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with few labels, and adapts when LLM information is imperfect, supported by a real‑data study of over 410K judgments from 21 LLM judges.

By Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
arXiv AI
Jun 3

Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge

arXiv:2512. 03019v2 Announce Type: replace-cross Abstract: Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consistency, or instruction-based self-aggregation) are inconsistent when ties are allowed.

By Hamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska, Mehdi Mirzazadeh
arXiv Machine Learning
Jun 8

Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies

arXiv:2606. 07492v1 Announce Type: cross Abstract: The ranking of recommendation algorithms is a challenging problem since model performance is sensitive to dataset characteristics such as sparsity, sequential structure, and scale.

By Ekaterina Grishina, Stepan Kuznetsov, Askar Tsyganov, Ilya Ivanov, Daria Korovaitceva, Margarita Rusanova, Uliana Parkina, Alexander Derevyagin, Evgeny Frolov, Sergey Samsonov, Anton Lysenko