arXiv:2511. 14117v2 Announce Type: replace Abstract: Supervised classifiers output a distribution over classes but are typically trained against a single label obtained by collapsing multiple annotators into a majority vote.
By Agamdeep Singh, Ashish Tiwari, Hosein Hasanbeig, Priyanshu Gupta
arXiv:2606. 05376v1 Announce Type: new Abstract: Many human-centered tasks, including natural language inference (NLI) and emotion recognition (ER), have multiple plausible interpretations, leading to label ambiguity and challenging disagreements across human annotators.
By Jingyao Wu, Ashley Wang, Keane Ong, Paul Pu Liang, Rosalind Picard
arXiv:2607. 18336v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance.
By Zilong Huang, Kong Aik Lee, Junjie Li, Zhe Li, Man-Wai Mak
arXiv:2607. 08493v1 Announce Type: new Abstract: Subjective NLP tasks often exhibit systematic annotator disagreement, requiring models that represent uncertainty rather than collapse it.
By Xia Cui, Ziyi Huang, N. R. Abeynayake
arXiv:2606. 00851v1 Announce Type: cross Abstract: Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues.
By Sukru Samet Dindar, Riki Shimizu, Xilin Jiang, Nima Mesgarani
arXiv:2608. 15619v1 Announce Type: new Abstract: Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipline.
By Keito Inoshita
arXiv:2608. 09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.
By Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols
arXiv:2606. 28772v1 Announce Type: cross Abstract: Hate speech annotation pipelines routinely collapse annotator disagreement into majority vote labels before training.
By Joshua Muhumuza, Joab Ezra Agaba, Mercy Amiyo
arXiv:2606. 16505v1 Announce Type: cross Abstract: Understanding speaker confidence is crucial in educational settings, as it can enhance personalised feedback and improve learning outcomes.
By Adam Wynn, Jingyun Wang, Xiangyu Tan
arXiv:2601. 12494v3 Announce Type: replace-cross Abstract: Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging.
By Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv:2606. 24941v2 Announce Type: replace-cross Abstract: Reviewing recorded interviews for affective cues such as composure and agitation is slow and subjective, and cloud services that could automate the task require sensitive audio to leave the device.
By Wai Laam Mak, Isibor Kennedy Ihianle, Pedro Machado
arXiv:2607. 21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks.
By Daniyal Kabir Dar, Arun Ross