Speech Emotion Recognition using Attention-based LSTM-Network with Residual Connection
arXiv:2606. 03359v1 Announce Type: cross Abstract: Speech emotion recognition is an important component of modern human-computer interaction systems.
Speech emotion recognition is an important component of modern human-computer interaction systems. However, many state-of-the-art approaches rely on large pretrained models with high computational and memory requirements, limiting their applicability.
arXiv:2606. 03359v1 Announce Type: cross Abstract: Speech emotion recognition is an important component of modern human-computer interaction systems.
arXiv:2606. 10278v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) aims to identify a speaker's emotional state from audio signals.
arXiv:2606. 08573v1 Announce Type: new Abstract: Speech emotion recognition (SER) is commonly formulated as utterance-level classification, although conversational emotion depends on a speaker's usual vocal range and the emotional context established by previous utterances.
arXiv:2607. 16803v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) is an important component in a wide range of human-centered applications, including healthcare, customer service, and human-omputer interaction.
arXiv:2606. 11197v1 Announce Type: cross Abstract: Speech-based automatic estimation of depression levels is essential for enabling early detection and timely intervention, particularly in resource-constrained mental health settings.
arXiv:2608. 05165v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data.
arXiv:2507. 07046v3 Announce Type: replace-cross Abstract: Nowadays, speech emotion recognition (SER) plays a vital role in the field of human-computer interaction (HCI) and the evolution of artificial intelligence (AI).
SISER is a speaker‑invariant speech emotion recognition framework that combines wav2vec 2.0 for feature extraction with an ECAPA‑TDNN speaker discriminator in an entropy‑based adversarial training scheme. By leveraging self‑supervised representations, SISER reduces reliance on large labeled datasets and suppresses speaker identity more effectively than shallow classifiers. On the IEMOCAP benchmark, SISER achieves a UA of 60.63%, surpassing both the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%).
DiaRelay introduces a lightweight adapter that lets large language models maintain a constant‑size dialogue‑level memory for emotion recognition in conversation. It builds on LoRA by adding a Selective Relay Memory Transition that aggregates useful historical evidence into a bounded memory, and a Dual‑axis Relay Memory Read that uses this memory to modulate low‑rank feature transformations. Experiments show DiaRelay achieves state‑of‑the‑art weighted F1 and accuracy on MELD with only 7.1 M additional trainable parameters, while also performing competitively on IEMOCAP.
The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.
The paper introduces ACERT, a module that incorporates a flexible-length window of conversational context to enhance Speech Emotion Recognition (SER). By capturing emotional evolution across utterances, ACERT outperforms state‑of‑the‑art methods on IEMOCAP, sets a new context‑aware benchmark on SAFE, and achieves strong results on MELD. Ablation studies attribute ACERT’s improvements to emotional and conversational continuity rather than speaker identity or acoustic conditions.
arXiv:2503. 00340v2 Announce Type: cross Abstract: Lightweight models are essential for real-time speech enhancement applications.