arXiv:2502.14643v3 Announce Type: replace
Abstract: Direct Preference Optimization (DPO) is a widely adopted offline algorithm for preference-based reinforcement learning from human feedback (RLHF),...
By Gengxu Li, Tingyu Xia, Yi Chang, Yuan Wu
arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.
By Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald, Katherine Metcalf
arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.
By Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu
arXiv:2609.21525v1 Announce Type: new
Abstract: Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being...
By Xuening Wu, Yanlan Kang, Shenqin Yin
arXiv:2609.38647v1 Announce Type: new
Abstract: Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting f...
By Sadegh Khorasani, Petrus Mikkola, Matthias Grossglauser
arXiv:2602. 09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences.
By Masanari Oi, Mahiro Ukai, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue
arXiv:2609.38860v1 Announce Type: cross
Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference lear...
By Zhongman Du, Huiming Zhang, Haodong Zhu, Baochang Zhang
arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.
By Hyung Gyu Rho
arXiv:2607. 25136v1 Announce Type: new Abstract: Research on preference optimization often varies the training objective while holding the data fixed.
By Zhengtao Yao, Runhao Li, Xupeng Chen, Jiayi Cheng, Chenqian Le, Michael Yue, Siheng Wang, Haoyan Xu, Yuqi Li, Chenhao Wei, Zhengdao Li, Rongchao Zhang, Guang Yang, Yidong Wang, Junhao Dong
arXiv:2605.02626v2 Announce Type: replace
Abstract: Direct Preference Optimization (DPO) improves relative preference by increasing the margin between chosen and rejected responses, but this objectiv...
By Inoussa Mouiche, Abderaouf Bahi
arXiv:2509. 22851v4 Announce Type: replace-cross Abstract: Margin-based optimization is fundamental to improving generalization and robustness in classification tasks.
By Yaswanth Chittepu, Prasann Singhal, Greg Durrett, Scott Niekum
arXiv:2607. 26094v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models.
By Yunpeng Chu