arXiv:2608.28932v1 Announce Type: new
Abstract: Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-o...
By Models Luc Debaupte, Tyler Baumgartner, Brandon Tai, Candice Fan, Bill Wang, Yi Zhong
VoiceLongMemEval (VLME) is a new benchmark that tests AI assistants on their ability to remember how users sounded by incorporating paralinguistic metadata—such as emotion labels, prosody descriptors, and voice events—into each conversational turn. The benchmark uses a three‑stage adversarial gate to ensure that models cannot succeed with transcript alone, revealing a significant affect gap: models gain 0.09 to 0.38 accuracy when provided with paralinguistic cues, and audio‑native models outperform standard ASR pipelines in extracting these signals. The dataset and code will be released upon acceptance.
By Ramit Pahwa, Parivesh Priye, Apoorva Beedu
The paper introduces a two‑stage speech anonymization framework that preserves both linguistic content and acoustic identity. It replaces personally identifiable information using a generative editing model and applies a flow‑matching anonymization technique (F3‑VA) to create diverse, distinct anonymized speakers. The authors evaluate privacy with speaker verification metrics and utility by training ASR, TTS, and SER models from scratch, showing stronger privacy protection with minimal utility loss compared to existing baselines.
By Yunchong Xiao, Yuxiang Zhao, Ziyang Ma, Shuai Wang, Kai Yu, Jiachun Liao, Xie Chen
arXiv:2507. 07046v3 Announce Type: replace-cross Abstract: Nowadays, speech emotion recognition (SER) plays a vital role in the field of human-computer interaction (HCI) and the evolution of artificial intelligence (AI).
By Shahana Yasmin Chowdhury, Bithi Banik, Md Tamjidul Hoque, Shreya Banerjee
arXiv:2609.05806v1 Announce Type: new
Abstract: Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide...
By Amir Ben Khalifa, Fanny Bezancon, Amine Trabelsi, Bessam Abdulrazak
Emotion recognition in conversation (ERC) is a production capability behind agent-assist prompts, escalation routing, and post-call analytics in contact-center-as-a-service (CCaaS) platforms, where co...
The paper introduces VoxPrivacy, a benchmark for assessing interactional privacy in Speech Language Models (SLMs). It evaluates models on a 32‑hour bilingual dataset across three difficulty tiers, revealing that most open‑source SLMs perform near random on conditional privacy decisions and even strong closed‑source systems struggle with proactive privacy inference. The authors also validate these findings on a real‑speech subset and show that fine‑tuning on a 4,000‑hour training set can improve privacy‑preserving capabilities while maintaining robustness.
By Yuxiang Wang, Hongyu Liu, Dekun Chen, Xueyao Zhang, Zhizheng Wu
The paper introduces an LLM-based framework for continuous dimensional emotion evaluation in multimodal dialogue, combining discrete emotion recognition with Valence-Arousal-Dominance (VAD) assessment on the IEMOCAP dataset. It incorporates acoustic cues as natural language descriptions via the SpeechCueLLM approach and evaluates six models from the LLaMA, GPT, and Qwen families using zero-shot, few-shot, and LoRA fine-tuning. LoRA-fine-tuned LLaMA models outperform prompt-engineered GPT models, achieving a new state-of-the-art Valence CCC of 0.7822, and ablation studies show that textual audio descriptions significantly benefit smaller models.
"whyItMatters":"The study demonstrates that domain adaptation through fine-tuning can surpass larger GPT models in multimodal emotion evaluation, highlighting the importance of tailored training for emotion recognition tasks."
By Yutong Hu, Jinho Choi
The paper evaluates three approaches for emotion recognition in conversation— a low‑cost stacked ensemble, an off‑the‑shelf LLM prompt, and a confidence‑gated hybrid that escalates only uncertain ensemble predictions to the LLM. Across three datasets (IEMOCAP, MELD, CMU‑MOSI), the hybrid consistently outperforms each pure system, achieving higher weighted F1 scores while routing most traffic through the inexpensive ensemble. This results in significant cost savings (≈$10‑85 per million utterances) and provides an interpretable escalation signal tied to emotion or sentiment shifts.
By Sai Babu Udayagiri, Arjun Chouhan, Ravisekhar Kanagala, Trishala Pavagada
arXiv:2606. 27717v1 Announce Type: cross Abstract: Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speech.
By Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang, Haonan Chen, Zeyu Jin
arXiv:2609.39453v1 Announce Type: cross
Abstract: Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent app...
By Hezhao Zhang, Thomas Hain
arXiv:2606. 00851v1 Announce Type: cross Abstract: Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues.
By Sukru Samet Dindar, Riki Shimizu, Xilin Jiang, Nima Mesgarani