arXiv Machine Learning By Dennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan Feuerriegel

Nonparametric LLM Evaluation from Preference Data

Read the original on arXiv Machine Learning →

arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment

The paper introduces MoPLEx, an expectation‑maximization algorithm for learning mixtures of Plackett‑Luce models from multi‑way ranking data. It augments rankings with synthetic responses from a base language model and uses a gradient‑based estimation to reduce inference cost, enabling efficient fitting of large‑scale models. Experiments show the method achieves low probability estimation error and improves clustering and ranking accuracy by 43.7% and 15.2% over baselines.

By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang