Autoregressive Direct Preference Optimization
arXiv:2602. 09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences.
The paper introduces LSC-DPO, a variant of Direct Preference Optimization that dynamically controls the learning signal to maintain sensitivity during training. By analyzing the logistic DPO loss geometrically, the authors identify the sigmoid factor as a key learning signal and propose a log‑space framework for stable target‑regime tracking. Experiments on AlpacaEval 2, MT‑Bench, and Anthropic‑HH demonstrate that LSC‑DPO outperforms standard DPO and other preference‑optimization baselines, and a signal‑budget compensation rule further reduces variability across different coefficient initializations.
arXiv:2602. 09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences.
arXiv:2502.14643v3 Announce Type: replace Abstract: Direct Preference Optimization (DPO) is a widely adopted offline algorithm for preference-based reinforcement learning from human feedback (RLHF),...
arXiv:2606. 19744v1 Announce Type: cross Abstract: Aligning language models with human preferences often requires optimising multiple behavioural objectives.
arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.
arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.
The paper investigates the Direct Preference Optimization (DPO) objective used for aligning language models, revealing that its coefficient β simultaneously controls both the inverse preference-noise scale and the optimization dynamics. This entanglement causes non‑monotonic policy deviation with respect to β and makes loss values incomparable across different β settings. The authors propose a centered‑softplus reformulation that decouples these effects, allowing independent tuning of the noise scale and learning‑rate, and provides a smooth β←0 limit that reduces to a linear preference‑margin objective.
arXiv:2606. 12505v1 Announce Type: cross Abstract: Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset.
The paper introduces Geometric Anchor Preference Optimization (GAPO), a method that replaces the static reference policy in Direct Preference Optimization with a dynamic, geometry-aware anchor—a small adversarial perturbation of the current policy. GAPO uses this anchor to adaptively reweight preference pairs based on local sensitivity, and defines an Anchor Gap that approximates worst‑case local margin degradation. Experiments show that GAPO improves robustness to noisy supervision while matching or surpassing existing LLM alignment and reasoning benchmarks.
arXiv:2607. 26094v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models.
The paper introduces GAW-PO, a gradient‑aligned token reweighting technique for Direct Preference Optimization (DPO). It assigns weaker penalties to rejected tokens whose gradients align with preferred behavior, while maintaining stronger penalties for conflicting tokens. Across 11 benchmarks in mathematics, reasoning, coding, and question answering, GAW‑PO outperforms standard DPO and other baselines, and remains robust when the DPO regularization parameter is reduced.
arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.
arXiv:2605.02626v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) improves relative preference by increasing the margin between chosen and rejected responses, but this objectiv...