You Can't Prefer Emotions You Don't Sample: Intensity Undershoot in DPO-Tuned LLMs
Read the original on arXiv AI →The study shows that instruction‑tuned language models, when asked to generate responses with varying emotional intensity, consistently undershoot the requested affect. By conditioning a model on continuous Valence‑Arousal targets and measuring output with a frozen regressor, the authors find that the gain for valence is only 0.26 and for arousal 0.13 on Llama‑3.1‑8B, far below the ideal value of 1. They trace this undershoot to the preference‑learning pipeline: training data such as EmoBank are neutral‑heavy and the candidate pool rarely contains extreme affect, so Direct Preference Optimization lacks examples to learn from. Expanding the target space uniformly and sampling a hotter, larger candidate pool raises valence gain to 0.40 and improves extrapolation with minimal in‑distribution cost, a result that also holds for Qwen3‑8B. Arousal remains more variable because the base model rarely generates highly aroused candidates.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.