arXiv Machine Learning

MORE-PLR: multi-output regression employed for partial label ranking

The paper introduces MORE-PLR, a method that tackles the partial label ranking problem by employing multi-output regression. It uses an encoder to transform incomplete rankings with ties into regression targets during training, and applies post‑hoc layers during inference to convert regression outputs into bucket orders. Experiments show that this framework competes with state‑of‑the‑art partial label ranking methods.

arXiv Machine Learning
Aug 27

Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment

The paper introduces MoPLEx, an expectation‑maximization algorithm for learning mixtures of Plackett‑Luce models from multi‑way ranking data. It augments rankings with synthetic responses from a base language model and uses a gradient‑based estimation to reduce inference cost, enabling efficient fitting of large‑scale models. Experiments show the method achieves low probability estimation error and improves clustering and ranking accuracy by 43.7% and 15.2% over baselines.

By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang