Hugging Face Blog

Preference Tuning LLMs with Direct Preference Optimization Methods

arXiv AI
Sep 17

A Zeroth-Order Paradigm for LLM Preference Alignment

The paper introduces Comparison-based Preference Optimization (ComPO), a zeroth-order method that aligns large language models with human preferences using comparison oracles instead of direct differentiable loss optimization. It provides theoretical convergence guarantees for both offline and online variants under smoothness, gradient sparsity, and oracle compatibility assumptions, and establishes performance bounds under local coverage and in-distribution reward accuracy. Experiments on several LLMs (Mistral, Llama, Gemma-2, Qwen3, Gemma-3) show that ComPO outperforms existing direct alignment methods, achieving higher length-controlled win rates and diagnostics that suggest mitigation of likelihood displacement.

By Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin
arXiv Computation and Language
Aug 28

Instruction Quality Matters: Refining Instructions for Effective Preference Learning

The paper investigates how the quality of instructions used to generate response pairs affects preference learning for language models. It shows that low‑quality or ambiguous instructions limit the range of response quality, weakening preference signals, and introduces an instruction‑refinement pipeline that improves data quality without discarding examples. Experiments across models and benchmarks demonstrate that refining instructions leads to better alignment and complements other data‑improvement methods.

By Seohyeong Lee, Hwaran Lee, Buru Chang
arXiv AI
Aug 24

VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models

The paper introduces VA‑DPO, a method that trains language models to generate text with a specified continuous affect point in the Valence‑Arousal plane. By using a frozen VA regressor to score candidate generations and selecting pairs with a distance margin, the approach modifies Direct Preference Optimization to better hit target emotions. Experiments on Llama‑3.1‑8B‑Instruct show a 33% reduction in mean VA distance compared to system‑prompting and 25% over few‑shot prompting, while maintaining performance on benchmarks like MMLU, HellaSwag, and TruthfulQA.

By Hyunwoo Kim