arXiv Machine Learning

Learning from Annotation Uncertainty: Entropy-Aware Curriculum for Speech Emotion Recognition

arXiv:2606. 27536v1 Announce Type: cross Abstract: Speech emotion recognition (SER) often relies on hard consensus labels that collapse annotator disagreement.

arXiv Computation and Language
Sep 4

SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training

SISER is a speaker‑invariant speech emotion recognition framework that combines wav2vec 2.0 for feature extraction with an ECAPA‑TDNN speaker discriminator in an entropy‑based adversarial training scheme. By leveraging self‑supervised representations, SISER reduces reliance on large labeled datasets and suppresses speaker identity more effectively than shallow classifiers. On the IEMOCAP benchmark, SISER achieves a UA of 60.63%, surpassing both the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%).

By Eunseo Choi, Hyunku Kang, Chanwoo Kim
arXiv Computation and Language
Sep 2

Post-hoc Alignment of LLM-judges to Human Judgment Distribution

The paper introduces NAPHA, a lightweight post‑hoc alignment method that improves large language model (LLM) predictions of human judgment distributions (HJD) by matching LLM output distributions to HJD through entropy‑based class assignment and specialized alignment models. Experiments on five datasets show that while LLMs perform near human‑level on hard‑label tasks, they struggle with soft‑label predictions, and NAPHA consistently enhances soft‑label accuracy, especially on high‑entropy instances. The study also demonstrates that better entropy class prediction can further boost NAPHA’s effectiveness.

By Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani
arXiv Computer Vision
Sep 16

Predicting Human Disagreement for Calibrated Dynamic Facial Expression Recognition

The paper introduces a disagreement‑aware dynamic facial expression recognition framework that directly learns from raw annotator vote vectors using a Dirichlet‑Multinomial likelihood, preserving both predictive mean and scale‑sensitive supervision. It adds an ambiguity head to estimate annotation entropy for unseen clips and employs a Chow‑style reject rule that integrates ambiguity, vacuity, temporal instability, and input quality for selective prediction. On the DFEW benchmark, the method maintains recognition accuracy while cutting expected calibration error by 30 % and area‑under‑risk‑curve by 15 %, with predicted ambiguity correlating 0.52 (Spearman) with true annotation entropy, and these gains transfer to FERV39k and hold under identity‑ and movie‑disjoint splits.

By Yiming Wang, Frederick W. B. Li, Jingyun Wang
arXiv Computation and Language
Sep 18

The Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation

The Public Discourse Corpus (PDC) is the first dataset of public‑figure interview speech annotated for affective valence and epistemic modality. It contains 998 videos from 100 speakers across seven professional domains, yielding 186,642 sentences (3.1 million words). A key methodological contribution is Target Speaker Participation (TSP), a five‑category annotation taxonomy with documented inter‑annotator reliability (κ = 0.616), and an audio‑first diarization pipeline that separates target‑speaker turns from interviewer and third‑party speech. The corpus, annotation tools, validation sample, and processing pipeline are released as open source.

By Bo Chen