Towards Robust Arabic Speech Emotion Recognition with Deep Learning
arXiv:2606. 10278v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) aims to identify a speaker's emotional state from audio signals.
arXiv:2607. 16803v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) is an important component in a wide range of human-centered applications, including healthcare, customer service, and human-omputer interaction.
arXiv:2606. 10278v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) aims to identify a speaker's emotional state from audio signals.
The paper evaluates deep learning models for electrocardiogram‑based emotion recognition, focusing on generalization across datasets rather than dataset‑specific performance. It introduces two open‑source tools—ARRC for standardized benchmarking and ARDT for inter‑dataset training—to merge three public AER datasets (CUADS, ASCERTAIN, DREAMER) into a more variable benchmark. Using these tools, the authors compare three prominent deep learning architectures and two CNN baselines with hyperparameter tuning and 10‑fold cross‑validation, revealing trade‑offs between accuracy and model complexity and providing a reproducible benchmark for future research.
arXiv:2606. 03359v1 Announce Type: cross Abstract: Speech emotion recognition is an important component of modern human-computer interaction systems.
Speech emotion recognition is an important component of modern human-computer interaction systems. However, many state-of-the-art approaches rely on large pretrained models with high computational and memory requirements, limiting their applicability.
SISER is a speaker‑invariant speech emotion recognition framework that combines wav2vec 2.0 for feature extraction with an ECAPA‑TDNN speaker discriminator in an entropy‑based adversarial training scheme. By leveraging self‑supervised representations, SISER reduces reliance on large labeled datasets and suppresses speaker identity more effectively than shallow classifiers. On the IEMOCAP benchmark, SISER achieves a UA of 60.63%, surpassing both the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%).
arXiv:2606. 11197v1 Announce Type: cross Abstract: Speech-based automatic estimation of depression levels is essential for enabling early detection and timely intervention, particularly in resource-constrained mental health settings.
arXiv:2507. 07046v3 Announce Type: replace-cross Abstract: Nowadays, speech emotion recognition (SER) plays a vital role in the field of human-computer interaction (HCI) and the evolution of artificial intelligence (AI).
The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.
The paper introduces ACERT, a module that incorporates a flexible-length window of conversational context to enhance Speech Emotion Recognition (SER). By capturing emotional evolution across utterances, ACERT outperforms state‑of‑the‑art methods on IEMOCAP, sets a new context‑aware benchmark on SAFE, and achieves strong results on MELD. Ablation studies attribute ACERT’s improvements to emotional and conversational continuity rather than speaker identity or acoustic conditions.
arXiv:2606. 25606v1 Announce Type: cross Abstract: Given the widespread prevalence of depression and its consequential impact on individuals and society, it is crucial to obtain objective measures for early diagnosis and intervention.
Given the widespread prevalence of depression and its consequential impact on individuals and society, it is crucial to obtain objective measures for early diagnosis and intervention. As a multidisciplinary topic, these objective measures should be interpretable and accessible to health care professionals, ensuring effective collaboration and treatment planning in the realm of mental health care.
arXiv:2608.28932v1 Announce Type: new Abstract: Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-o...