arXiv AI By Hyunwoo Kim, Usama Khalid

You Can't Prefer Emotions You Don't Sample: Intensity Undershoot in DPO-Tuned LLMs

Read the original on arXiv AI →

The study shows that instruction‑tuned language models, when asked to generate responses with varying emotional intensity, consistently undershoot the requested affect. By conditioning a model on continuous Valence‑Arousal targets and measuring output with a frozen regressor, the authors find that the gain for valence is only 0.26 and for arousal 0.13 on Llama‑3.1‑8B, far below the ideal value of 1. They trace this undershoot to the preference‑learning pipeline: training data such as EmoBank are neutral‑heavy and the candidate pool rarely contains extreme affect, so Direct Preference Optimization lacks examples to learn from. Expanding the target space uniformly and sampling a hotter, larger candidate pool raises valence gain to 0.40 and improves extrapolation with minimal in‑distribution cost, a result that also holds for Qwen3‑8B. Arousal remains more variable because the base model rarely generates highly aroused candidates.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 24

VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models

The paper introduces VA‑DPO, a method that trains language models to generate text with a specified continuous affect point in the Valence‑Arousal plane. By using a frozen VA regressor to score candidate generations and selecting pairs with a distance margin, the approach modifies Direct Preference Optimization to better hit target emotions. Experiments on Llama‑3.1‑8B‑Instruct show a 33% reduction in mean VA distance compared to system‑prompting and 25% over few‑shot prompting, while maintaining performance on benchmarks like MMLU, HellaSwag, and TruthfulQA.

By Hyunwoo Kim
arXiv AI
Jun 29

When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model

arXiv:2606. 21641v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been proposed as hyperparameter-optimization (HPO) advisors that "warm-start" search from prior knowledge, proposing strong configurations in very few evaluations.

By Carson Rodrigues, Oysturn Vas, Isaiah Abner DCosta, Nithish Kumar Prabhakaran
arXiv AI
Jul 7

Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning

arXiv:2607. 04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorly understood.

By Faid Keddouri, Sohaib Houhou, Aissa Boulmerka, Nadir Farhi