arXiv AI

Affective Flow Language Model for Emotional Support Conversation

The paper introduces the Affective Flow Language Model (AFlow), which treats multi‑turn emotional support conversations as an evolving affective utility flow along dialogue trajectories. AFlow searches diverse support paths, estimates utilities of intermediate states, and employs Affective Flow Preference Optimization (AFPO) to propagate downstream preference signals to earlier states, enabling consistent strategy transitions. Experiments on ExTES and ESConv demonstrate improved strategy alignment, response diversity, and generation quality across various model settings.

arXiv AI
Jul 31

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

arXiv:2603. 06194v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for large language models (LLMs) has shown strong performance in single-turn tasks, but extending it to multi-turn interaction remains challenging due to sparse rewards and poor per-turn credit assignment.

By Naifan Zhang, Ruihan Sun, Jinwei Su, Hengjie Yang, Zhengyuan Pan, Zhaohan Chen, Xiaofan Zhang
arXiv Computation and Language
Sep 3

PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue Generation

PragAlign is a feedback‑guided framework that improves synthetic dialogue generation by iteratively generating, evaluating, and revising conversations to meet specified service context, target intent, and target emotion. Using an LLM‑based evaluator that scores intent alignment, emotion alignment, coherence, fluency, and overall quality, PragAlign achieves a 99.50% acceptance rate on 800 dialogue specifications, outperforming one‑shot and repeated generation without feedback. Human evaluation confirms that intent expression and dialogue flow are reliably recognized, while emotion appropriateness remains more variable.

By Smitha Muthya Sudheendra, Jaideep Srivastava
arXiv AI
3d ago

Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue

The paper introduces an LLM-based framework for continuous dimensional emotion evaluation in multimodal dialogue, combining discrete emotion recognition with Valence-Arousal-Dominance (VAD) assessment on the IEMOCAP dataset. It incorporates acoustic cues as natural language descriptions via the SpeechCueLLM approach and evaluates six models from the LLaMA, GPT, and Qwen families using zero-shot, few-shot, and LoRA fine-tuning. LoRA-fine-tuned LLaMA models outperform prompt-engineered GPT models, achieving a new state-of-the-art Valence CCC of 0.7822, and ablation studies show that textual audio descriptions significantly benefit smaller models. "whyItMatters":"The study demonstrates that domain adaptation through fine-tuning can surpass larger GPT models in multimodal emotion evaluation, highlighting the importance of tailored training for emotion recognition tasks."

By Yutong Hu, Jinho Choi
arXiv AI
Aug 19

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.

By Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury