arXiv:2502.14643v3 Announce Type: replace
Abstract: Direct Preference Optimization (DPO) is a widely adopted offline algorithm for preference-based reinforcement learning from human feedback (RLHF),...
By Gengxu Li, Tingyu Xia, Yi Chang, Yuan Wu
The paper investigates the Direct Preference Optimization (DPO) objective used for aligning language models, revealing that its coefficient β simultaneously controls both the inverse preference-noise scale and the optimization dynamics. This entanglement causes non‑monotonic policy deviation with respect to β and makes loss values incomparable across different β settings. The authors propose a centered‑softplus reformulation that decouples these effects, allowing independent tuning of the noise scale and learning‑rate, and provides a smooth β←0 limit that reduces to a linear preference‑margin objective.
By Ivan Kruzhilov
arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.
By Hyung Gyu Rho
arXiv:2604. 18239v4 Announce Type: replace-cross Abstract: Preference optimization is widely used to align large language models (LLMs) with human preferences.
By Wei Chen, Yubing Wu, Junmei Yang, Delu Zeng, Qibin Zhao, John Paisley, Min Chen, Zhou Wang
arXiv:2606. 12505v1 Announce Type: cross Abstract: Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset.
By Pengwei Sun
The paper introduces GAP-DPO, a method for personalizing large language models by selecting preference pairs based on gradient alignment with user utility. It formalizes personalized preference learning as a geometry‑aligned optimization problem, showing that off‑policy sampling can shift DPO updates from error correction to reinforcement when preference margins align with utility gradients. Experiments demonstrate that GAP‑DPO improves stylistic fidelity, preference alignment, and overall generation quality over standard DPO variants.
By Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou
arXiv:2602. 09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences.
By Masanari Oi, Mahiro Ukai, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue
arXiv:2606. 19744v1 Announce Type: cross Abstract: Aligning language models with human preferences often requires optimising multiple behavioural objectives.
By Pranav Bhandari, Nicolas Fay, Amitava Datta, Usman Naseem, Mehwish Nasim
arXiv:2509. 03647v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models.
By Dani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas, Simon Fu, Narmeen Oozeer
arXiv:2510. 04773v2 Announce Type: replace Abstract: As Large Language Models (LLMs) demonstrate remarkable capabilities learned from vast corpora, concerns regarding data privacy and safety are receiving increasing attention.
By Kai Qin, Jiaqi Wu, Jianxiang He, Haoyuan Sun, Yifei Zhao, Xu Wang, Bin Liang, Yongzhe Chang, Cheng Li, Tiantian Zhang, Houde Liu
GroupDPO introduces a memory‑efficient approach to group‑wise direct preference optimization for aligning large language models. By using first‑order linearization with per‑response coefficients, the method decouples samples during backpropagation, dramatically reducing peak memory usage and enabling scalable training with larger groups. Experiments in both offline and online settings show that leveraging multiple responses consistently outperforms single‑pair training, and adding a negative log‑likelihood term on positive responses is essential for performance gains and training stability.
By Jixuan Leng, Si Si, Hsiang-Fu Yu, Vinod Raman, Inderjit S. Dhillon
arXiv:2603.13359v2 Announce Type: replace
Abstract: Language models are commonly fine-tuned for safety alignment to refuse harmful prompts. One approach fine-tunes them to generate categorical refusa...
By Rishab Alagharu, Ishneet Sukhvinder Singh, Shaibi Shamsudeen, Zhen Wu, Ashwinee Panda