arXiv:2511. 14117v2 Announce Type: replace Abstract: Supervised classifiers output a distribution over classes but are typically trained against a single label obtained by collapsing multiple annotators into a majority vote.
By Agamdeep Singh, Ashish Tiwari, Hosein Hasanbeig, Priyanshu Gupta
arXiv:2608.28932v1 Announce Type: new
Abstract: Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-o...
By Models Luc Debaupte, Tyler Baumgartner, Brandon Tai, Candice Fan, Bill Wang, Yi Zhong
arXiv:2609.05806v1 Announce Type: new
Abstract: Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide...
By Amir Ben Khalifa, Fanny Bezancon, Amine Trabelsi, Bessam Abdulrazak
arXiv:2606. 05376v1 Announce Type: new Abstract: Many human-centered tasks, including natural language inference (NLI) and emotion recognition (ER), have multiple plausible interpretations, leading to label ambiguity and challenging disagreements across human annotators.
By Jingyao Wu, Ashley Wang, Keane Ong, Paul Pu Liang, Rosalind Picard
arXiv:2607. 18336v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance.
By Zilong Huang, Kong Aik Lee, Junjie Li, Zhe Li, Man-Wai Mak
SISER is a speaker‑invariant speech emotion recognition framework that combines wav2vec 2.0 for feature extraction with an ECAPA‑TDNN speaker discriminator in an entropy‑based adversarial training scheme. By leveraging self‑supervised representations, SISER reduces reliance on large labeled datasets and suppresses speaker identity more effectively than shallow classifiers. On the IEMOCAP benchmark, SISER achieves a UA of 60.63%, surpassing both the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%).
By Eunseo Choi, Hyunku Kang, Chanwoo Kim
arXiv:2607. 08493v1 Announce Type: new Abstract: Subjective NLP tasks often exhibit systematic annotator disagreement, requiring models that represent uncertainty rather than collapse it.
By Xia Cui, Ziyi Huang, N. R. Abeynayake
arXiv:2609.39453v1 Announce Type: cross
Abstract: Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent app...
By Hezhao Zhang, Thomas Hain
arXiv:2606. 00851v1 Announce Type: cross Abstract: Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues.
By Sukru Samet Dindar, Riki Shimizu, Xilin Jiang, Nima Mesgarani
The paper introduces NAPHA, a lightweight post‑hoc alignment method that improves large language model (LLM) predictions of human judgment distributions (HJD) by matching LLM output distributions to HJD through entropy‑based class assignment and specialized alignment models. Experiments on five datasets show that while LLMs perform near human‑level on hard‑label tasks, they struggle with soft‑label predictions, and NAPHA consistently enhances soft‑label accuracy, especially on high‑entropy instances. The study also demonstrates that better entropy class prediction can further boost NAPHA’s effectiveness.
By Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani
The paper introduces a disagreement‑aware dynamic facial expression recognition framework that directly learns from raw annotator vote vectors using a Dirichlet‑Multinomial likelihood, preserving both predictive mean and scale‑sensitive supervision. It adds an ambiguity head to estimate annotation entropy for unseen clips and employs a Chow‑style reject rule that integrates ambiguity, vacuity, temporal instability, and input quality for selective prediction. On the DFEW benchmark, the method maintains recognition accuracy while cutting expected calibration error by 30 % and area‑under‑risk‑curve by 15 %, with predicted ambiguity correlating 0.52 (Spearman) with true annotation entropy, and these gains transfer to FERV39k and hold under identity‑ and movie‑disjoint splits.
By Yiming Wang, Frederick W. B. Li, Jingyun Wang
The Public Discourse Corpus (PDC) is the first dataset of public‑figure interview speech annotated for affective valence and epistemic modality. It contains 998 videos from 100 speakers across seven professional domains, yielding 186,642 sentences (3.1 million words). A key methodological contribution is Target Speaker Participation (TSP), a five‑category annotation taxonomy with documented inter‑annotator reliability (κ = 0.616), and an audio‑first diarization pipeline that separates target‑speaker turns from interviewer and third‑party speech. The corpus, annotation tools, validation sample, and processing pipeline are released as open source.
By Bo Chen