A Better Spur Should Start From Each Objective
arXiv:2609.08211v1 Announce Type: new Abstract: Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing...
arXiv:2505. 10892v2 Announce Type: replace Abstract: Post-training LLMs with RLHF and preference optimization methods (e.
arXiv:2609.08211v1 Announce Type: new Abstract: Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing...
The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.
arXiv:2512. 15765v3 Announce Type: replace Abstract: Data valuation is a natural framework for understanding which preference datasets matter most when aligning a Large Language Model (LLM) using multiple sources.
The paper introduces Comparison-based Preference Optimization (ComPO), a zeroth-order method that aligns large language models with human preferences using comparison oracles instead of direct differentiable loss optimization. It provides theoretical convergence guarantees for both offline and online variants under smoothness, gradient sparsity, and oracle compatibility assumptions, and establishes performance bounds under local coverage and in-distribution reward accuracy. Experiments on several LLMs (Mistral, Llama, Gemma-2, Qwen3, Gemma-3) show that ComPO outperforms existing direct alignment methods, achieving higher length-controlled win rates and diagnostics that suggest mitigation of likelihood displacement.
The paper introduces GAP-DPO, a method for personalizing large language models by selecting preference pairs based on gradient alignment with user utility. It formalizes personalized preference learning as a geometry‑aligned optimization problem, showing that off‑policy sampling can shift DPO updates from error correction to reinforcement when preference margins align with utility gradients. Experiments demonstrate that GAP‑DPO improves stylistic fidelity, preference alignment, and overall generation quality over standard DPO variants.
arXiv:2504.20106v4 Announce Type: replace Abstract: Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessiv...
arXiv:2506. 13702v4 Announce Type: replace-cross Abstract: Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative to pairwise preference learning by directly leveraging scalar feedback.
arXiv:2604.13175v2 Announce Type: replace-cross Abstract: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While sing...
The paper investigates how to create steerable AI models that can balance multiple, sometimes conflicting objectives, a necessity for pluralistic alignment. Using Multi-Objective Direct Preference Optimization (MODPO), the authors examine when a single model can improve two objectives simultaneously and how to cover many trade‑offs without training separate models. They find that two pre‑training measurements predict objective alignment for human‑annotated data but not for AI‑annotated data, and that selecting the nearest trained model or merging parameters can broaden trade‑off coverage, though neither approach consistently matches direct training.
The paper introduces Suan, a new preference optimization algorithm designed to improve safety alignment in large language models. Suan operates directly at the gradient level, avoiding traditional variational derivations, which yields more interpretable and robust training dynamics. Experiments show that Suan outperforms existing methods, achieving superior safety alignment while maintaining response utility.
Aligning large language models with diverse and heterogeneous human values requires multi-objective alignment methods to effectively trade off conflicting preference dimensions. Current methods achieve this trade-off by training policies conditioned on preference vectors and leveraging online direct preference optimization.
arXiv:2510. 01167v2 Announce Type: replace-cross Abstract: Aligning large language models to human preferences is inherently multidimensional, yet most pipelines collapse heterogeneous signals into a single objective.