The paper introduces GAP-DPO, a method for personalizing large language models by selecting preference pairs based on gradient alignment with user utility. It formalizes personalized preference learning as a geometry‑aligned optimization problem, showing that off‑policy sampling can shift DPO updates from error correction to reinforcement when preference margins align with utility gradients. Experiments demonstrate that GAP‑DPO improves stylistic fidelity, preference alignment, and overall generation quality over standard DPO variants.
By Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou
arXiv:2505. 10892v2 Announce Type: replace Abstract: Post-training LLMs with RLHF and preference optimization methods (e.
By Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng Wen
arXiv:2504.20106v4 Announce Type: replace
Abstract: Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessiv...
By Ren-Wei Liang, Chin-Ting Hsu, Chan-Hung Yu, Saransh Agrawal, Shih-Cheng Huang, Chieh-Yen Lin, Shang-Tse Chen, Kuan-Hao Huang, Shao-Hua Sun
The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.
By Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi, Alireza Hashemi, Pouya M. Ghari, Joseph Hoover, James Rae, Morteza Dehghani
arXiv:2609.08211v1 Announce Type: new
Abstract: Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing...
By Shanwen Mao, Hao Zhang, Guangtao nie, Zhiheng Li, Huimu Wang, Sulong Xu, Gu Simiu
UniPolicy is a unified objective‑specific policy framework for search advertising that jointly optimizes relevance, click propensity, and commercial value. It uses objective‑aware prefix tokens, sparse MoE‑LoRA routing, and residual FFNs to decouple parameters within a shared backbone, and constructs pairwise preferences from multi‑stage behavioral feedback to strengthen clicked candidates. In large‑scale offline tests and a 7‑day online A/B test, UniPolicy improves CTR by 0.71%, RPS by 1.58%, and advertising revenue by 1.32% while keeping serving latency stable.
By Kun Yao, Yuhang Zhou, Yichi Zhang, Zeliang Tong, Shengri Xue, Haitao Wang, Siyu Lu, Qianlong Xie, Xingxing Wang
MiCRo is a two‑stage framework that improves personalized preference learning for large language models. It first uses a context‑aware mixture model to capture diverse human preferences from large binary preference datasets, then applies an online routing strategy to dynamically adjust mixture weights based on context, reducing ambiguity. Experiments on multiple datasets show that MiCRo captures diverse preferences and enhances downstream personalization.
By Jingyan Shen, Jiarui Yao, Rui Yang, Yifan Sun, Feng Luo, Rui Pan, Tong Zhang, Han Zhao
The paper introduces MoPLEx, an expectation‑maximization algorithm for learning mixtures of Plackett‑Luce models from multi‑way ranking data. It augments rankings with synthetic responses from a base language model and uses a gradient‑based estimation to reduce inference cost, enabling efficient fitting of large‑scale models. Experiments show the method achieves low probability estimation error and improves clustering and ranking accuracy by 43.7% and 15.2% over baselines.
By Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang
UniPolicy is a multi-policy alignment framework for search advertising that jointly optimizes relevance, click propensity, and commercial value. It uses objective-specific prefix tokens, sparse MoE-LoRA routing, and residual FFNs to decouple parameters within a shared backbone, and builds pairwise preferences from multi-stage behavioral feedback to improve generation. In large-scale offline tests and a 7‑day online A/B test, UniPolicy achieved balanced gains across metrics, boosting CTR by 0.71%, RPS by 1.58%, and revenue by 1.32% while keeping latency stable.
The paper investigates how to create steerable AI models that can balance multiple, sometimes conflicting objectives, a necessity for pluralistic alignment. Using Multi-Objective Direct Preference Optimization (MODPO), the authors examine when a single model can improve two objectives simultaneously and how to cover many trade‑offs without training separate models. They find that two pre‑training measurements predict objective alignment for human‑annotated data but not for AI‑annotated data, and that selecting the nearest trained model or merging parameters can broaden trade‑off coverage, though neither approach consistently matches direct training.
By David Tsoi, Esra D\"onmez
arXiv:2509. 22851v4 Announce Type: replace-cross Abstract: Margin-based optimization is fundamental to improving generalization and robustness in classification tasks.
By Yaswanth Chittepu, Prasann Singhal, Greg Durrett, Scott Niekum
arXiv:2606. 04284v1 Announce Type: cross Abstract: Preference modeling plays a central role in reinforcement learning from human feedback (RLHF), enabling large language models (LLMs) to align with human values.
By Yifan Wang, Jinyi Mu, Mayank Jobanputra, Yu Wang, Ji-Ung Lee, Soyoung Oh, Isabel Valera, Vera Demberg