arXiv Machine Learning By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang

Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment

Read the original on arXiv Machine Learning →

The paper introduces MoPLEx, an expectation‑maximization algorithm for learning mixtures of Plackett‑Luce models from multi‑way ranking data. It augments rankings with synthetic responses from a base language model and uses a gradient‑based estimation to reduce inference cost, enabling efficient fitting of large‑scale models. Experiments show the method achieves low probability estimation error and improves clustering and ranking accuracy by 43.7% and 15.2% over baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.