arXiv AI By Yang Qu, Yusheng Han, Chengjia Feng, Handan Liu

LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization

Read the original on arXiv AI →

The paper introduces LSC-DPO, a variant of Direct Preference Optimization that dynamically controls the learning signal to maintain sensitivity during training. By analyzing the logistic DPO loss geometrically, the authors identify the sigmoid factor as a key learning signal and propose a log‑space framework for stable target‑regime tracking. Experiments on AlpacaEval 2, MT‑Bench, and Anthropic‑HH demonstrate that LSC‑DPO outperforms standard DPO and other preference‑optimization baselines, and a signal‑budget compensation rule further reduces variability across different coefficient initializations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 10

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.

By Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu
arXiv Machine Learning
Aug 28

Disentangling Optimization Scale from Preference Scale in DPO

The paper investigates the Direct Preference Optimization (DPO) objective used for aligning language models, revealing that its coefficient β simultaneously controls both the inverse preference-noise scale and the optimization dynamics. This entanglement causes non‑monotonic policy deviation with respect to β and makes loss values incomparable across different β settings. The authors propose a centered‑softplus reformulation that decouples these effects, allowing independent tuning of the noise scale and learning‑rate, and provides a smooth β←0 limit that reduces to a linear preference‑margin objective.

By Ivan Kruzhilov