arXiv:2608.28932v1 Announce Type: new
Abstract: Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-o...
By Models Luc Debaupte, Tyler Baumgartner, Brandon Tai, Candice Fan, Bill Wang, Yi Zhong
The paper introduces a discriminative adaptation for SpeechLLMs that reads the hidden state of the final prompt token via a simple classification head, enabling emotion recognition in a single forward pass without altering the backbone. This approach replaces the generative decoder, which can produce out‑of‑set labels and favor frequent classes, with a controlled comparison between generative and discriminative inference. Experiments on IEMOCAP show improved Macro F1 scores, elimination of hallucinations, and larger gains on realistic ASR transcripts, while revealing that emotion directions encode indirect associations reflecting web‑scale text biases.
By Hasindri Watawana, Sergio Burdisso, Esa\'u Villatoro-Tello, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke
arXiv:2512.07571v3 Announce Type: replace
Abstract: This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tune...
By Nicolas Calbucura, Jose Guillen, Valentin Barriere
arXiv:2609.39453v1 Announce Type: cross
Abstract: Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent app...
By Hezhao Zhang, Thomas Hain
arXiv:2606. 00670v1 Announce Type: cross Abstract: Face-to-face speech comprehension is inherently multimodal, integrating acoustic signals with visible articulation, facial expression, head motion, and other socially relevant cues.
By Zhou Yang, Yueyi Yang
The paper introduces ACERT, a module that incorporates a flexible-length window of conversational context to enhance Speech Emotion Recognition (SER). By capturing emotional evolution across utterances, ACERT outperforms state‑of‑the‑art methods on IEMOCAP, sets a new context‑aware benchmark on SAFE, and achieves strong results on MELD. Ablation studies attribute ACERT’s improvements to emotional and conversational continuity rather than speaker identity or acoustic conditions.
By Arthur Peuvot, Romaric Besan\c{c}on, Ga\"el de Chalendar, Bianca Vieru, Ioana Vasilescu
arXiv:2609.05871v1 Announce Type: cross
Abstract: Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-s...
By Song-ha Jo, Sehyun Lee, Soyoon Kim, Jaesik Choi, Sanghyuk Choi
arXiv:2608. 06409v1 Announce Type: cross Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation.
By Linkai Peng, Baorian Nuchged
arXiv:2607. 15755v1 Announce Type: cross Abstract: Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions.
By Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu, Haizhou Li
VoiceLongMemEval (VLME) is a new benchmark that tests AI assistants on their ability to remember how users sounded by incorporating paralinguistic metadata—such as emotion labels, prosody descriptors, and voice events—into each conversational turn. The benchmark uses a three‑stage adversarial gate to ensure that models cannot succeed with transcript alone, revealing a significant affect gap: models gain 0.09 to 0.38 accuracy when provided with paralinguistic cues, and audio‑native models outperform standard ASR pipelines in extracting these signals. The dataset and code will be released upon acceptance.
By Ramit Pahwa, Parivesh Priye, Apoorva Beedu
Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion Recognition (SER) remains a significant bottleneck. Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate in flat Euclidean space, and this geometry fails to capture the multi-granularity nature of emotion cues, which range from low-level prosody to high-level semantics.
The paper introduces an LLM-based framework for continuous dimensional emotion evaluation in multimodal dialogue, combining discrete emotion recognition with Valence-Arousal-Dominance (VAD) assessment on the IEMOCAP dataset. It incorporates acoustic cues as natural language descriptions via the SpeechCueLLM approach and evaluates six models from the LLaMA, GPT, and Qwen families using zero-shot, few-shot, and LoRA fine-tuning. LoRA-fine-tuned LLaMA models outperform prompt-engineered GPT models, achieving a new state-of-the-art Valence CCC of 0.7822, and ablation studies show that textual audio descriptions significantly benefit smaller models.
"whyItMatters":"The study demonstrates that domain adaptation through fine-tuning can surpass larger GPT models in multimodal emotion evaluation, highlighting the importance of tailored training for emotion recognition tasks."
By Yutong Hu, Jinho Choi