arXiv Machine Learning By Zahra Omidi, John H. L. Hansen

Learning from Annotation Uncertainty: Entropy-Aware Curriculum for Speech Emotion Recognition

Read the original on arXiv Machine Learning →

arXiv:2606. 27536v1 Announce Type: cross Abstract: Speech emotion recognition (SER) often relies on hard consensus labels that collapse annotator disagreement.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 4

SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training

SISER is a speaker‑invariant speech emotion recognition framework that combines wav2vec 2.0 for feature extraction with an ECAPA‑TDNN speaker discriminator in an entropy‑based adversarial training scheme. By leveraging self‑supervised representations, SISER reduces reliance on large labeled datasets and suppresses speaker identity more effectively than shallow classifiers. On the IEMOCAP benchmark, SISER achieves a UA of 60.63%, surpassing both the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%).

By Eunseo Choi, Hyunku Kang, Chanwoo Kim