arXiv AI By Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Read the original on arXiv AI →

arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.