The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.
By Xu Lin, Ke Wang, Hui Kang, Xinying Wang
arXiv:2606. 15038v1 Announce Type: new Abstract: Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift.
By Zhemin Zhang, Weijie Chen, David Le, Amara Tariq, Alex Wallace, Matthew Stib, Juan Maria Farina, Chadi Ayoub, Reza Arsanjani, Imon Banerjee
The paper introduces RAFM-SER++, a lightweight multimodal speech emotion recognition framework designed for real‑time surveillance systems. It replaces heavy bidirectional cross‑modal transformers with an asymmetric Residual Attention Fusion Mechanism that injects affective speech cues into text representations via a one‑directional residual attention pathway. Experiments on IEMOCAP and ESD show RAFM‑SER++ outperforms the HuBERT‑Base baseline and MemoCMT, reducing trainable parameters by over 60%, achieving 79.60 it/s inference speed, and reaching BACC scores of 81.10% on IEMOCAP and 95.39% on ESD.
By Ngo Truong Dinh, Tung-Lam Bui, Chi-Trung Duong, Vien Nguyen Thi, Viet-Anh Nguyen, Phuc-Lu Le
RAMamba-Net is a new multimodal fusion network designed for auditory attention decoding (AAD) that combines EEG and electrooculography (EOG) signals. It uses a Mamba-enhanced band-aware convolutional Transformer to capture EEG band-specific patterns and long-range temporal dynamics, while a dual-branch encoder models EOG temporal and inter-channel dependencies. Cross‑modal attention and a reliability‑aware module estimate sample‑wise modality weights, improving fusion robustness and achieving a 5.76% accuracy gain over unimodal baselines on two AAD benchmarks.
By Xingyi He, Ziwei Wang, Dongrui Wu
arXiv:2504. 03707v2 Announce Type: replace-cross Abstract: Emotion recognition is crucial for advancing mental health, healthcare, and technologies such as brain-computer interfaces.
By Md Niaz Imtiaz, Naimul Khan
arXiv:2607. 00358v1 Announce Type: new Abstract: Electroencephalogram (EEG) captures endogenous brain activity with high temporal fidelity and holds substantial promise for precise emotion decoding.
By Xin Zhou, Xiang Zhang, Hao Deng, Lijun Yin
The paper introduces CoMA-DiT, a bidirectional cross‑modal Diffusion Transformer that uses paired modalities as mutual generative supervision for latent augmentation rather than just inputs for fusion. By conditioning velocity prediction on the paired modality through cross‑modal attention and injecting variation via a reliability‑gated residual mechanism, CoMA‑DiT improves multimodal brain state decoding. Experiments on auditory attention decoding and emotion recognition show consistent gains over 20 baselines, with absolute accuracy and macro‑F1 improvements of 4.28% and 6.70% respectively, and extensive analyses confirm its robustness and interpretability.
By Ziwei Wang, Xingyi He, Hongbin Wang, Tianwang Jia, Bohan Fang, Dongrui Wu
arXiv:2601. 07565v2 Announce Type: replace-cross Abstract: Multimodal emotion understanding requires the integration of heterogeneous data sources, including text, audio, and visual modalities, while simultaneously addressing discrete emotion recognition and continuous sentiment analysis.
By Jiaqi Qiao, Xinran Li, Yifan Lyu, Xiujuan Xu, Liu Yu
arXiv:2607. 04139v1 Announce Type: new Abstract: Self-supervised learning (SSL) shows strong potential for cross-dataset transfer by improving feature representation and generalization.
By Huqin Weng, Jiayang Huang, Yimin Wen, Jie Du, Chi-Man Vong, Chuangquan Chen
Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations.
Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition proposes a hybrid framework that combines a Transformer and a Graph Attention Network to capture both global semantic information and fine-grained relationships between modalities. The model is evaluated on the IEMOCAP and MELD datasets, achieving weighted F1 scores of 72.45% and 77.37%, respectively, and surpasses state‑of‑the‑art methods. These results suggest that integrating multimodal features with balanced global and local context modeling can provide deeper emotional insights for dialogue emotion recognition.
By Jiaqi Qiao, Yifan Lyu, Xiujuan Xu
The paper introduces ‘One Model for All’, a universal pre‑training framework that tackles EEG‑based emotion recognition across diverse datasets and paradigms. It decouples learning into a univariate self‑supervised contrastive pre‑training stage using a Unified Channel Schema, followed by a multivariate fine‑tuning stage that employs an Adaptive Resampling Transformer and a Graph Attention Network to model spatio‑temporal dependencies. Experiments demonstrate state‑of‑the‑art performance on within‑subject benchmarks (SEED 99.27%, DEAP 93.69%, DREAMER 93.93%) and superior cross‑dataset transfer, with ablation studies highlighting the critical role of the GAT module.
By Xiang Li, You Li, Yazhou Zhang