arXiv Machine Learning

AFA-Net: A Differential Attention Approach for Auditory Attention Detection

AFA‑Net introduces a differential attention mechanism to improve Auditory Attention Detection (AAD) from EEG signals, explicitly targeting noisy data. The framework outperforms existing deep learning models, achieving 96.8% accuracy within a 2‑second decision window while using fewer parameters. It represents one of the first approaches to actively mitigate EEG noise in AAD tasks.

arXiv AI
Sep 12

RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection

RAMamba-Net is a new multimodal fusion network designed for auditory attention decoding (AAD) that combines EEG and electrooculography (EOG) signals. It uses a Mamba-enhanced band-aware convolutional Transformer to capture EEG band-specific patterns and long-range temporal dynamics, while a dual-branch encoder models EOG temporal and inter-channel dependencies. Cross‑modal attention and a reliability‑aware module estimate sample‑wise modality weights, improving fusion robustness and achieving a 5.76% accuracy gain over unimodal baselines on two AAD benchmarks.

By Xingyi He, Ziwei Wang, Dongrui Wu
arXiv AI
Sep 10

Adaptive Anisotropic Attention for Axis-Structured Signals

The paper introduces Adaptive Anisotropic Attention (AAA), a method that splits self‑attention into temporal and spatial paths for axis‑structured signals like EEG. A learned gate combines the two paths for each token, and the resulting AXON model outperforms dense attention baselines on six EEG tasks and shows transfer to audio spectrograms. The study demonstrates that aligning attention with the natural axes of structured data provides a beneficial inductive bias.

By Mahir Jain, Parshva Runwal, Aditya Ray Mishra, Arvasu Kulkarni, Sandeep Singh, Siddharth Panwar
Hugging Face Trending Papers
Sep 8

Adaptive Anisotropic Attention for Axis-Structured Signals

The paper introduces Adaptive Anisotropic Attention (AAA), a method that splits self‑attention into temporal and spatial paths for axis‑structured signals like EEG. A learned gate combines the two paths for each token, and the resulting AXON model outperforms dense attention baselines on six EEG tasks and shows benefits in audio spectrogram experiments. The study demonstrates that aligning attention with natural signal axes provides a useful inductive bias.

arXiv Machine Learning
3d ago

Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition

The paper introduces EmoSpeechBrain, a multimodal emotion recognition framework that fuses EEG and speech signals. It employs differential attention in the EEG encoder to cancel shared noise and an attention-based gating adapter to align modalities and weight their contributions. Experiments on PME4 and EAV datasets show up to 12.9% accuracy improvement over other EEG encoders and surpass unimodal baselines by up to 23.1%.

By Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen
arXiv AI
Sep 25

SHINE: Sequential Hierarchical Integration Network for EEG and MEG

SHINE is a Sequential Hierarchical Integration Network designed to reconstruct speech envelope and Mel spectrogram from EEG and MEG recordings. It uses a residual sensor adapter, dilated-block states for temporal depth, and a target- and time-dependent gate to fuse hierarchical and attention-enhanced context predictions. Across two EEG and two MEG datasets, SHINE achieved the highest mean envelope and mean-Mel Pearson correlations among nine baseline methods and ranked second in the NeurIPS 2025 PNPL Competition’s speech-detection Extended Track.

By Xiran Xu, Yujie Yan, Songyi Li, Linze Zheng, Zifeng Zhang, Mochu Dong, Jing Chen
Hugging Face Trending Papers
Jun 23

Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training

Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance can be driven by trial-specific EEG structure that acts as shortcuts for target selection, leading to poor generalization on unseen trials.