Hugging Face Blog

Direct Preference Optimization Beyond Chatbots

arXiv AI
Jun 10

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.

By Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu
arXiv AI
Sep 3

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy

The paper surveys how AI copilots—AI-powered assistants for knowledge workers and developers—can personalize their behavior by optimizing user preferences. It reviews how preference signals are collected, modeled at different interaction stages, and refined through feedback loops, and introduces a taxonomy of optimization techniques for pre-, mid-, and post-interaction phases. The study evaluates each technique’s strengths, limitations, and design implications, aiming to unify efforts across AI personalization, human‑AI interaction, and language model adaptation.

By Saleh Afzoon, Ali Shahsavandi, Phuong Thao Huynh, Melika Zare, Zahra Jahanandish, Amin Beheshti, Usman Naseem