The study shows that instruction‑tuned language models, when asked to generate responses with varying emotional intensity, consistently undershoot the requested affect. By conditioning a model on continuous Valence‑Arousal targets and measuring output with a frozen regressor, the authors find that the gain for valence is only 0.26 and for arousal 0.13 on Llama‑3.1‑8B, far below the ideal value of 1. They trace this undershoot to the preference‑learning pipeline: training data such as EmoBank are neutral‑heavy and the candidate pool rarely contains extreme affect, so Direct Preference Optimization lacks examples to learn from. Expanding the target space uniformly and sampling a hotter, larger candidate pool raises valence gain to 0.40 and improves extrapolation with minimal in‑distribution cost, a result that also holds for Qwen3‑8B. Arousal remains more variable because the base model rarely generates highly aroused candidates.
By Hyunwoo Kim, Usama Khalid
arXiv:2606. 19744v1 Announce Type: cross Abstract: Aligning language models with human preferences often requires optimising multiple behavioural objectives.
By Pranav Bhandari, Nicolas Fay, Amitava Datta, Usman Naseem, Mehwish Nasim
arXiv:2606. 12505v1 Announce Type: cross Abstract: Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset.
By Pengwei Sun
arXiv:2602. 09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences.
By Masanari Oi, Mahiro Ukai, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue
arXiv:2607. 25136v1 Announce Type: new Abstract: Research on preference optimization often varies the training objective while holding the data fixed.
By Zhengtao Yao, Runhao Li, Xupeng Chen, Jiayi Cheng, Chenqian Le, Michael Yue, Siheng Wang, Haoyan Xu, Yuqi Li, Chenhao Wei, Zhengdao Li, Rongchao Zhang, Guang Yang, Yidong Wang, Junhao Dong
arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.
By Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald, Katherine Metcalf