arXiv Computation and Language By Eunseo Choi, Hyunku Kang, Chanwoo Kim

SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training

Read the original on arXiv Computation and Language →

SISER is a speaker‑invariant speech emotion recognition framework that combines wav2vec 2.0 for feature extraction with an ECAPA‑TDNN speaker discriminator in an entropy‑based adversarial training scheme. By leveraging self‑supervised representations, SISER reduces reliance on large labeled datasets and suppresses speaker identity more effectively than shallow classifiers. On the IEMOCAP benchmark, SISER achieves a UA of 60.63%, surpassing both the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 18

Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation

arXiv:2603. 10827v2 Announce Type: replace-cross Abstract: Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or specific fields such as emotions or the speaker's gender, leaving it unclear whether they encode speaker identity.

By Thomas Thebaud, Yuzhe Wang, Laureano Moro-Velazquez, Jesus Villalba-Lopez, Najim Dehak
arXiv AI
Jul 14

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

arXiv:2607. 09767v1 Announce Type: cross Abstract: The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech.

By Adrien Schneider (M-PSI), Kacper Zabkowski (M-PSI), Anderson Augusma (M-PSI), Fr\'ed\'erique Letu\'e (SAM, SVH), Maria Camila Pinzon (M-PSI), Dominique Vaufreydaz (M-PSI)
arXiv Computation and Language
Sep 23

Enriching Speech Emotion Representations with Conversational Context

The paper introduces ACERT, a module that incorporates a flexible-length window of conversational context to enhance Speech Emotion Recognition (SER). By capturing emotional evolution across utterances, ACERT outperforms state‑of‑the‑art methods on IEMOCAP, sets a new context‑aware benchmark on SAFE, and achieves strong results on MELD. Ablation studies attribute ACERT’s improvements to emotional and conversational continuity rather than speaker identity or acoustic conditions.

By Arthur Peuvot, Romaric Besan\c{c}on, Ga\"el de Chalendar, Bianca Vieru, Ioana Vasilescu