Normalized Rewards for Preference Optimization
arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.
arXiv:2608. 14828v1 Announce Type: new Abstract: Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no real help.
arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.
arXiv:2604. 18239v4 Announce Type: replace-cross Abstract: Preference optimization is widely used to align large language models (LLMs) with human preferences.
The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.
arXiv:2606. 00291v1 Announce Type: cross Abstract: In RLHF, each training example contains a prompt $x$ and two candidate responses $y,y'$, and annotators provide pairwise preferences between these responses.
Swiss-Knife is a framework that extends decode‑time alignment for frozen language models by treating the alignment specification as a runtime object. It introduces hot‑swappable scoring blades, a batch normaliser, a pairwise aggregation operator, and a selection rule, and characterises admissible aggregation operators with a representation theorem. In experiments, Swiss‑Knife paired with DPO‑LoRA blades and an uncertainty‑aware pairwise tournament outperforms six existing decode‑time methods, achieving a higher harmonic F1 score, lower refusal rate, and faster objective reconfiguration.
arXiv:2606. 16771v1 Announce Type: new Abstract: As LLMs advance, post-training reinforcement learning (RL) increasingly relies on multi-dimensional rewards to cultivate comprehensive capabilities.
arXiv:2509. 23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences.
arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.
arXiv:2608. 10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity.
arXiv:2605.02626v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) improves relative preference by increasing the margin between chosen and rejected responses, but this objectiv...
arXiv:2607. 02755v1 Announce Type: cross Abstract: Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned.
arXiv:2607. 04590v1 Announce Type: new Abstract: Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences.