arXiv AI

LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization

The paper introduces LSC-DPO, a variant of Direct Preference Optimization that dynamically controls the learning signal to maintain sensitivity during training. By analyzing the logistic DPO loss geometrically, the authors identify the sigmoid factor as a key learning signal and propose a log‑space framework for stable target‑regime tracking. Experiments on AlpacaEval 2, MT‑Bench, and Anthropic‑HH demonstrate that LSC‑DPO outperforms standard DPO and other preference‑optimization baselines, and a signal‑budget compensation rule further reduces variability across different coefficient initializations.

arXiv AI
Jun 10

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.

By Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu
arXiv Machine Learning
Aug 28

Disentangling Optimization Scale from Preference Scale in DPO

The paper investigates the Direct Preference Optimization (DPO) objective used for aligning language models, revealing that its coefficient β simultaneously controls both the inverse preference-noise scale and the optimization dynamics. This entanglement causes non‑monotonic policy deviation with respect to β and makes loss values incomparable across different β settings. The authors propose a centered‑softplus reformulation that decouples these effects, allowing independent tuning of the noise scale and learning‑rate, and provides a smooth β←0 limit that reduces to a linear preference‑margin objective.

By Ivan Kruzhilov
arXiv Machine Learning
Sep 29

Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment

The paper introduces Geometric Anchor Preference Optimization (GAPO), a method that replaces the static reference policy in Direct Preference Optimization with a dynamic, geometry-aware anchor—a small adversarial perturbation of the current policy. GAPO uses this anchor to adaptively reweight preference pairs based on local sensitivity, and defines an Anchor Gap that approximates worst‑case local margin degradation. Experiments show that GAPO improves robustness to noisy supervision while matching or surpassing existing LLM alignment and reasoning benchmarks.

By Youngjae Cho, Jongsuk Kim, Ji-Hoon Kim
arXiv Computation and Language
4d ago

GAW-PO: Preference Optimization with Gradient-Aligned Token Weights

The paper introduces GAW-PO, a gradient‑aligned token reweighting technique for Direct Preference Optimization (DPO). It assigns weaker penalties to rejected tokens whose gradients align with preferred behavior, while maintaining stronger penalties for conflicting tokens. Across 11 benchmarks in mathematics, reasoning, coding, and question answering, GAW‑PO outperforms standard DPO and other baselines, and remains robust when the DPO regularization parameter is reduced.

By Andreea Dutulescu, Stefan Ruseti, Mihai Masala, Traian Rebedea, Mihai Dascalu