Robust Nash Alignment introduces a game-theoretic framework that seeks a policy with a high worst-case win rate against both an adversarial competitor and any preference kernel within an ambiguity set around a nominal preference. The authors propose a four-player primal-dual proxy game and an optimistic mirror descent-ascent algorithm to efficiently optimize this robust objective, proving convergence guarantees and demonstrating improved performance in controlled tabular games and LLM alignment experiments.
By Shihab Ahmed, Debamita Ghosh, David Tang, Yudan Wang, Alvaro Velasquez, Yue Wang
arXiv:2608.25200v2 Announce Type: replace-cross
Abstract: We study learning a mixture of $k$ Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous und...
By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang
arXiv:2604. 27733v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human intent, whether through explicit reward modeling or direct methods such as DPO, fundamentally relies on minimizing a surrogate loss as a proxy for the true pairwise ranking objective.
By Mehryar Mohri, Yutao Zhong
The paper introduces Comparison-based Preference Optimization (ComPO), a zeroth-order method that aligns large language models with human preferences using comparison oracles instead of direct differentiable loss optimization. It provides theoretical convergence guarantees for both offline and online variants under smoothness, gradient sparsity, and oracle compatibility assumptions, and establishes performance bounds under local coverage and in-distribution reward accuracy. Experiments on several LLMs (Mistral, Llama, Gemma-2, Qwen3, Gemma-3) show that ComPO outperforms existing direct alignment methods, achieving higher length-controlled win rates and diagnostics that suggest mitigation of likelihood displacement.
By Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin
The paper introduces MoPLEx, an expectation‑maximization algorithm for learning mixtures of Plackett‑Luce models from multi‑way ranking data. It augments rankings with synthetic responses from a base language model and uses a gradient‑based estimation to reduce inference cost, enabling efficient fitting of large‑scale models. Experiments show the method achieves low probability estimation error and improves clustering and ranking accuracy by 43.7% and 15.2% over baselines.
By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang
arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.
By Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald, Katherine Metcalf