Direct Preference Optimization for Chatbot Fine-Tuning: An Empirical Study
arXiv:2606. 12881v2 Announce Type: replace-cross Abstract: We present an approach to fine-tuning large language models using Direct Preference Optimization (DPO), a reinforcement learning technique.
Related stories
A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.
Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
arXiv:2508. 11847v4 Announce Type: replace-cross Abstract: We propose a method for evaluating the robustness of widely used LLM ranking systems -- variants of a Bradley--Terry model -- to dropping a worst-case very small fraction of preference data.
Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebycheff Scalarization
arXiv:2604.13175v2 Announce Type: replace-cross Abstract: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While sing...
Optimal Design for Active Preference Learning with Biased LLM Judges
arXiv:2609.38860v1 Announce Type: cross Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference lear...
EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
arXiv:2603.02041v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are predominantly trained on English-centric data, resulting in uneven performance for smaller languages. We stu...
Length-Controlled Margin-Based Preference Optimization without Reference Model
arXiv:2502.14643v3 Announce Type: replace Abstract: Direct Preference Optimization (DPO) is a widely adopted offline algorithm for preference-based reinforcement learning from human feedback (RLHF),...
CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations
CroCo introduces cross‑lingual contrastive preference tuning on self‑generations, extending prior English‑only methods to 14 high‑ and low‑resource languages. A reward model trained solely on English preferences, applied to a multilingual base, yields effective within‑language rankings and improves performance in both monolingual and multilingual settings without catastrophic forgetting. The approach requires on‑policy data; off‑policy responses and online preference optimization offer limited gains, yet on structured tasks CroCo matches or surpasses the base model in most languages, and on open‑ended generation it wins 28/30 judge evaluations across 15 languages.
Efficiently Aligning Language Models with Online Natural Language Feedback
arXiv:2605. 04356v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards has been used to elicit impressive performance from language models in many domains.
Distributionally Robust Reinforcement Learning with Human Feedback
arXiv:2503. 00539v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs).
Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation
arXiv:2609.06893v1 Announce Type: cross Abstract: Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational eff...
Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
arXiv:2608. 15949v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and natural dialogue.