arXiv Computation and Language By Arthur Peuvot, Romaric Besan\c{c}on, Ga\"el de Chalendar, Bianca Vieru, Ioana Vasilescu

Enriching Speech Emotion Representations with Conversational Context

Read the original on arXiv Computation and Language →

The paper introduces ACERT, a module that incorporates a flexible-length window of conversational context to enhance Speech Emotion Recognition (SER). By capturing emotional evolution across utterances, ACERT outperforms state‑of‑the‑art methods on IEMOCAP, sets a new context‑aware benchmark on SAFE, and achieves strong results on MELD. Ablation studies attribute ACERT’s improvements to emotional and conversational continuity rather than speaker identity or acoustic conditions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 30

TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

arXiv:2606. 30543v1 Announce Type: cross Abstract: With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important.

By Sathvik Manikantan Napa Ugandhar, Hao Zhang, Alison Gunzler, Yuzhe Wang, Thomas Thebaud, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Vel\'azquez
arXiv Machine Learning
Sep 7

Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation

The paper introduces DSSM-CRF, an audio‑only architecture for conversational speech emotion recognition that separates cross‑speaker contextual influence from within‑speaker emotion evolution. It uses bidirectional state‑space models to encode fused self‑supervised speech representations at both frame and dialogue scales, then orders each speaker’s utterances into an independent dynamic conditional random field chain. The model achieves state‑of‑the‑art performance on IEMOCAP and MELD, with complementary gains from speaker‑wise factorization and CRF modeling.

By Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen
arXiv AI
Sep 2

VoiceLongMemEval: Do Assistants Remember How You Sounded?

VoiceLongMemEval (VLME) is a new benchmark that tests AI assistants on their ability to remember how users sounded by incorporating paralinguistic metadata—such as emotion labels, prosody descriptors, and voice events—into each conversational turn. The benchmark uses a three‑stage adversarial gate to ensure that models cannot succeed with transcript alone, revealing a significant affect gap: models gain 0.09 to 0.38 accuracy when provided with paralinguistic cues, and audio‑native models outperform standard ASR pipelines in extracting these signals. The dataset and code will be released upon acceptance.

By Ramit Pahwa, Parivesh Priye, Apoorva Beedu
arXiv Computation and Language
Sep 22

COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning

arXiv:2609.22697v1 Announce Type: new Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated...

By Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue, Boyi Kang, Shilei Zhang, Jinglei Liu, Yue Wang, Junlan Feng, Bei Liu, Wei Xue