MoPLEx: Estimating Plackett-Luce Mixture Models for Multi-Objective Alignment
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces MoPLEx, an expectation‑maximization algorithm for learning mixtures of Plackett‑Luce models from multi‑way ranking data. It augments rankings with synthetic responses from a base language model and uses a gradient‑based estimation to reduce inference cost, enabling efficient fitting of large‑scale models. Experiments show the method achieves low probability estimation error and improves clustering and ranking accuracy by 43.7% and 15.2% over baselines.
arXiv:2607. 01715v1 Announce Type: new Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level.
arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.
The paper introduces Comparison-based Preference Optimization (ComPO), a zeroth-order method that aligns large language models with human preferences using comparison oracles instead of direct differentiable loss optimization. It provides theoretical convergence guarantees for both offline and online variants under smoothness, gradient sparsity, and oracle compatibility assumptions, and establishes performance bounds under local coverage and in-distribution reward accuracy. Experiments on several LLMs (Mistral, Llama, Gemma-2, Qwen3, Gemma-3) show that ComPO outperforms existing direct alignment methods, achieving higher length-controlled win rates and diagnostics that suggest mitigation of likelihood displacement.
UniPolicy is a unified objective‑specific policy framework for search advertising that jointly optimizes relevance, click propensity, and commercial value. It uses objective‑aware prefix tokens, sparse MoE‑LoRA routing, and residual FFNs to decouple parameters within a shared backbone, and constructs pairwise preferences from multi‑stage behavioral feedback to strengthen clicked candidates. In large‑scale offline tests and a 7‑day online A/B test, UniPolicy improves CTR by 0.71%, RPS by 1.58%, and advertising revenue by 1.32% while keeping serving latency stable.
CORE improves compositional reasoning in multimodal language models by distilling a cross‑attentive reranker’s fine‑grained judgments into the embedding model. It generates candidate lists across five compositional matching levels and trains with a Rank‑KL objective to replicate the reranker’s ranking. Experiments on COLA, SUGARCREPE++, and NEGBENCH show CORE‑RERANKER‑8B outperforms Jina‑Reranker by 10.7 points, while CORE‑EMBED‑8B achieves the best overall average among evaluated embeddings, with gains also transferring to the MCMR benchmark without harming COCO or Flickr30K retrieval.