arXiv:2606. 00851v1 Announce Type: cross Abstract: Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues.
By Sukru Samet Dindar, Riki Shimizu, Xilin Jiang, Nima Mesgarani
arXiv:2609.05806v1 Announce Type: new
Abstract: Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide...
By Amir Ben Khalifa, Fanny Bezancon, Amine Trabelsi, Bessam Abdulrazak
arXiv:2607.23648v2 Announce Type: replace
Abstract: Using large language models (LLMs) to assist psychological counseling is an important task in the field of natural language processing. The constru...
By Kaitong Weng, Lixin Liu, Zihao Liu, Bo Wang, Shiguang Ni
The paper introduces EmoStance, a method for controlling the affective orientation of empathetic responses by leveraging weak supervision from emoji distributions. It builds the EmojiDialogue dataset, extending EmpatheticDialogues with emoji votes and confidence scores, and uses a frozen instruction‑tuned LLM steered by continuous prefix embeddings to generate responses that align with the listener’s stance. In blind pairwise evaluations, EmoStance achieves a 62.2% decisive win rate, notably improving contextual specificity and perceived responsiveness compared to baselines.
By Ziyuan Jin, Yuxuan Ge, Zheng Tian
The paper introduces EmoStance, a method for controlling the affective orientation of empathetic responses in dialogue systems. It leverages weak supervision from multi‑annotator emoji distributions to create a latent control space that approximates listener stance, and uses a frozen instruction‑tuned LLM steered by continuous prefix embeddings. Evaluation shows a 62.2% decisive win rate over baselines, especially in contextual specificity and perceived responsiveness.
arXiv:2606. 09837v1 Announce Type: cross Abstract: Emotional interaction is increasingly crucial for conversational AI, yet current systems lack a self-emotion determination mechanism to drive the streaming text-to-speech (TTS) synthesis.
By Yue Zhao, Hongyan Li, Yong Chen, Luo Ji
Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-c...
The paper introduces DSSM-CRF, an audio‑only architecture for conversational speech emotion recognition that separates cross‑speaker contextual influence from within‑speaker emotion evolution. It uses bidirectional state‑space models to encode fused self‑supervised speech representations at both frame and dialogue scales, then orders each speaker’s utterances into an independent dynamic conditional random field chain. The model achieves state‑of‑the‑art performance on IEMOCAP and MELD, with complementary gains from speaker‑wise factorization and CRF modeling.
By Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen
arXiv:2504. 11837v3 Announce Type: replace-cross Abstract: Emotional support conversation (ESC) aims to alleviate people's emotional distress through effective conversations.
By Yue Zhao, Qingqing Gu, Xiaoyu Wang, Teng Chen, Zhonglin Jiang, Yong Chen, Hongyan Li, Luo Ji
arXiv:2607. 15282v1 Announce Type: cross Abstract: Empathy is most often theorized as resonance: a mirroring of another's present emotional or cognitive state.
By Molood Arman
arXiv:2606. 30543v1 Announce Type: cross Abstract: With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important.
By Sathvik Manikantan Napa Ugandhar, Hao Zhang, Alison Gunzler, Yuzhe Wang, Thomas Thebaud, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Vel\'azquez
The paper investigates the role of minimal responses—short, empathic utterances—in psychological counseling, noting that such brief replies are common in human dialogues but underrepresented in large language model (LLM) outputs. Using a two‑stage filtering approach and contextual verification with an LLM, the authors systematically analyze minimal responses across multiple counseling datasets. They find that while strong commercial LLMs can produce minimal replies when prompted, they often fail to judge when these replies are appropriate, and counseling‑specific models trained on synthetic data tend to generate longer, content‑rich responses instead.
By Zhiyang Qi