arXiv AI

DynamicPO: Dynamic Preference Optimization for Recommendation

arXiv:2605. 00327v2 Announce Type: replace-cross Abstract: In large language model (LLM)-based recommendation systems, direct preference optimization (DPO) effectively aligns recommendations with user preferences, requiring multi-negative objective functions to leverage abundant implicit-feedback negatives and sharpen preference boundaries.

arXiv AI
Jun 10

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.

By Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu
arXiv Computation and Language
Sep 3

GroupDPO: Memory-Efficient Group-Wise Direct Preference Optimization

GroupDPO introduces a memory‑efficient approach to group‑wise direct preference optimization for aligning large language models. By using first‑order linearization with per‑response coefficients, the method decouples samples during backpropagation, dramatically reducing peak memory usage and enabling scalable training with larger groups. Experiments in both offline and online settings show that leveraging multiple responses consistently outperforms single‑pair training, and adding a negative log‑likelihood term on positive responses is essential for performance gains and training stability.

By Jixuan Leng, Si Si, Hsiang-Fu Yu, Vinod Raman, Inderjit S. Dhillon
arXiv AI
2d ago

Gradient-Aligned Pair Selection for Personalized Preference Optimization

The paper introduces GAP-DPO, a method for personalizing large language models by selecting preference pairs based on gradient alignment with user utility. It formalizes personalized preference learning as a geometry‑aligned optimization problem, showing that off‑policy sampling can shift DPO updates from error correction to reinforcement when preference margins align with utility gradients. Experiments demonstrate that GAP‑DPO improves stylistic fidelity, preference alignment, and overall generation quality over standard DPO variants.

By Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou
arXiv AI
Sep 24

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

COPE (Continual Optimization with Personalized embedding and self-Evaluation) is a new framework that continually personalizes large language models using learnable user embeddings and self‑evaluation to generate proxy rewards. It integrates preference capture, self‑evaluation calibration, and personalized response optimization into a single update step, allowing continuous model updates even when explicit user feedback is sparse. Experiments demonstrate that COPE outperforms both training‑free and training‑based baselines, remains complementary to Retrieval‑Augmented Prompting, and shows reliable self‑evaluation, meaningful preference patterns, stable general capabilities, and robustness to shifting preferences and alternative evaluators.

By Ruike Cao, Fugen Yao, Liang Dong, Jian Xu, Guanjun Jiang, Li Xiao
arXiv AI
Sep 25

DeGRe: Dense-supervised Generative Reranking for Recommendation

DeGRe is a dense‑supervised generative reranking framework designed to improve multi‑stage recommender systems by addressing label bias and credit assignment issues. It uses an offline Lookahead Evaluator with beam search to generate dense supervision signals, which are distilled into a lightweight Online Generator that can perform efficient greedy decoding at inference time. Experiments show that DeGRe outperforms baselines on public benchmarks and industrial datasets, and it has been successfully deployed on Taobao Flash Shopping to enhance online recommendations.

By Chaotian Song, Jingyao Zhang, Chenghao Chen, Zisen Sang, Dehai Zhao, Guodong Cao, Boxi Wu, Deng Cai, Jia Jia