arXiv Machine Learning By Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen

Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition

Read the original on arXiv Machine Learning →

The paper introduces EmoSpeechBrain, a multimodal emotion recognition framework that fuses EEG and speech signals. It employs differential attention in the EEG encoder to cancel shared noise and an attention-based gating adapter to align modalities and weight their contributions. Experiments on PME4 and EAV datasets show up to 12.9% accuracy improvement over other EEG encoders and surpass unimodal baselines by up to 23.1%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 2

UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment

arXiv:2606. 00170v1 Announce Type: cross Abstract: In recent years, emotion recognition based on physiological signals such as electroencephalogram (EEG) has gained considerable attention, as internal physiological data offer greater objectivity and reliability compared to external behavioral data like facial expressions.

By Zheng Wang, Shuo Wang, Junhong Wang
arXiv AI
Sep 12

Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding

The paper introduces CoMA-DiT, a bidirectional cross‑modal Diffusion Transformer that uses paired modalities as mutual generative supervision for latent augmentation rather than just inputs for fusion. By conditioning velocity prediction on the paired modality through cross‑modal attention and injecting variation via a reliability‑gated residual mechanism, CoMA‑DiT improves multimodal brain state decoding. Experiments on auditory attention decoding and emotion recognition show consistent gains over 20 baselines, with absolute accuracy and macro‑F1 improvements of 4.28% and 6.70% respectively, and extensive analyses confirm its robustness and interpretability.

By Ziwei Wang, Xingyi He, Hongbin Wang, Tianwang Jia, Bohan Fang, Dongrui Wu
arXiv AI
Sep 12

RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection

RAMamba-Net is a new multimodal fusion network designed for auditory attention decoding (AAD) that combines EEG and electrooculography (EOG) signals. It uses a Mamba-enhanced band-aware convolutional Transformer to capture EEG band-specific patterns and long-range temporal dynamics, while a dual-branch encoder models EOG temporal and inter-channel dependencies. Cross‑modal attention and a reliability‑aware module estimate sample‑wise modality weights, improving fusion robustness and achieving a 5.76% accuracy gain over unimodal baselines on two AAD benchmarks.

By Xingyi He, Ziwei Wang, Dongrui Wu
arXiv Computer Vision
Sep 4

Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG

The paper presents a method for emotion recognition in virtual reality where head‑mounted displays occlude the upper face. By fusing lower‑face video with electromyography (EMG) signals from the occluded upper face, the authors achieve a 51% macro‑F1 score across seven emotional categories, outperforming image‑only and EMG‑only baselines. A new synchronized multimodal dataset from 20 participants is introduced and will be shared under an ethical‑use agreement.

By Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel, Mustafa Tevfik Lafci, David Przewozny, Anna Hilsmann, Peter Eisert, Sebastian Bosse