A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition
arXiv:2608. 09088v1 Announce Type: new Abstract: Mixed emotions represent a clinically relevant but still underexplored target for automatic emotion recognition.
The paper introduces EmoSpeechBrain, a multimodal emotion recognition framework that fuses EEG and speech signals. It employs differential attention in the EEG encoder to cancel shared noise and an attention-based gating adapter to align modalities and weight their contributions. Experiments on PME4 and EAV datasets show up to 12.9% accuracy improvement over other EEG encoders and surpass unimodal baselines by up to 23.1%.
arXiv:2608. 09088v1 Announce Type: new Abstract: Mixed emotions represent a clinically relevant but still underexplored target for automatic emotion recognition.
arXiv:2608. 15999v1 Announce Type: new Abstract: Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, modality-specific feature-extraction pipelines before fusion.
arXiv:2606. 00170v1 Announce Type: cross Abstract: In recent years, emotion recognition based on physiological signals such as electroencephalogram (EEG) has gained considerable attention, as internal physiological data offer greater objectivity and reliability compared to external behavioral data like facial expressions.
The paper introduces CoMA-DiT, a bidirectional cross‑modal Diffusion Transformer that uses paired modalities as mutual generative supervision for latent augmentation rather than just inputs for fusion. By conditioning velocity prediction on the paired modality through cross‑modal attention and injecting variation via a reliability‑gated residual mechanism, CoMA‑DiT improves multimodal brain state decoding. Experiments on auditory attention decoding and emotion recognition show consistent gains over 20 baselines, with absolute accuracy and macro‑F1 improvements of 4.28% and 6.70% respectively, and extensive analyses confirm its robustness and interpretability.
RAMamba-Net is a new multimodal fusion network designed for auditory attention decoding (AAD) that combines EEG and electrooculography (EOG) signals. It uses a Mamba-enhanced band-aware convolutional Transformer to capture EEG band-specific patterns and long-range temporal dynamics, while a dual-branch encoder models EOG temporal and inter-channel dependencies. Cross‑modal attention and a reliability‑aware module estimate sample‑wise modality weights, improving fusion robustness and achieving a 5.76% accuracy gain over unimodal baselines on two AAD benchmarks.
The paper presents a method for emotion recognition in virtual reality where head‑mounted displays occlude the upper face. By fusing lower‑face video with electromyography (EMG) signals from the occluded upper face, the authors achieve a 51% macro‑F1 score across seven emotional categories, outperforming image‑only and EMG‑only baselines. A new synchronized multimodal dataset from 20 participants is introduced and will be shared under an ethical‑use agreement.
arXiv:2504. 03707v2 Announce Type: replace-cross Abstract: Emotion recognition is crucial for advancing mental health, healthcare, and technologies such as brain-computer interfaces.
The paper introduces RAFM-SER++, a lightweight multimodal speech emotion recognition framework designed for real‑time surveillance systems. It replaces heavy bidirectional cross‑modal transformers with an asymmetric Residual Attention Fusion Mechanism that injects affective speech cues into text representations via a one‑directional residual attention pathway. Experiments on IEMOCAP and ESD show RAFM‑SER++ outperforms the HuBERT‑Base baseline and MemoCMT, reducing trainable parameters by over 60%, achieving 79.60 it/s inference speed, and reaching BACC scores of 81.10% on IEMOCAP and 95.39% on ESD.
arXiv:2606. 10718v1 Announce Type: cross Abstract: Electroencephalography (EEG) is a widely adopted technique for monitoring brain activity, offering valuable insights into neurological states due to its high temporal resolution and cost-effectiveness.
arXiv:2607. 18336v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance.
The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.
arXiv:2607. 04139v1 Announce Type: new Abstract: Self-supervised learning (SSL) shows strong potential for cross-dataset transfer by improving feature representation and generalization.