arXiv AI

Subject-Invariant Cross-Modal Decoding of Perceived Speech from Brain Recordings

The paper introduces the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which fuses fMRI and MEG data to decode perceived speech from non‑invasive brain signals. Comprehensive experiments show that SICMD improves Top‑1, Top‑10, and Rankacc scores by over 10%, 10%, and 1.7% respectively, while cutting training costs by 88.8% and 60.5% compared to existing multi‑subject and intra‑subject approaches. Visualizations further confirm the method’s effectiveness.

arXiv AI
Aug 25

Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings

The paper introduces a Cross-Subject Perceived Speech Decoding (CPSD) framework that tackles the challenge of decoding perceived speech from non‑invasive brain recordings across different subjects. CPSD uses a two‑stage training process: first, contrastive learning pre‑trains a source model on multiple subjects to capture shared representations; second, personal specialization fine‑tunes the model for a target subject by extracting consistent components and further training on that subject’s data. A Positional Encoding‑based Spatial Attention (PESA) module is added to remap MEG/EEG data into a standardized reference space, improving cross‑subject consistency. Evaluations on three datasets (Armeni 2022, PKUEEG 2025, Broderick 2018) show that CPSD outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top‑10 accuracy, demonstrating its effectiveness, efficiency, and robustness.

By Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen
arXiv AI
Sep 12

Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding

The paper introduces CoMA-DiT, a bidirectional cross‑modal Diffusion Transformer that uses paired modalities as mutual generative supervision for latent augmentation rather than just inputs for fusion. By conditioning velocity prediction on the paired modality through cross‑modal attention and injecting variation via a reliability‑gated residual mechanism, CoMA‑DiT improves multimodal brain state decoding. Experiments on auditory attention decoding and emotion recognition show consistent gains over 20 baselines, with absolute accuracy and macro‑F1 improvements of 4.28% and 6.70% respectively, and extensive analyses confirm its robustness and interpretability.

By Ziwei Wang, Xingyi He, Hongbin Wang, Tianwang Jia, Bohan Fang, Dongrui Wu
arXiv AI
Sep 25

SHINE: Sequential Hierarchical Integration Network for EEG and MEG

SHINE is a Sequential Hierarchical Integration Network designed to reconstruct speech envelope and Mel spectrogram from EEG and MEG recordings. It uses a residual sensor adapter, dilated-block states for temporal depth, and a target- and time-dependent gate to fuse hierarchical and attention-enhanced context predictions. Across two EEG and two MEG datasets, SHINE achieved the highest mean envelope and mean-Mel Pearson correlations among nine baseline methods and ranked second in the NeurIPS 2025 PNPL Competition’s speech-detection Extended Track.

By Xiran Xu, Yujie Yan, Songyi Li, Linze Zheng, Zifeng Zhang, Mochu Dong, Jing Chen
arXiv AI
Sep 7

A Roadmap for MEG Foundation Models

The article outlines the emerging field of Magnetoencephalography (MEG) foundation models, explaining how these reusable, pretrained models can surpass traditional task‑specific decoding pipelines. It reviews current design choices—such as tokenization, sensor versus source representations, and self‑supervised objectives—and notes the limited number of existing MEG‑specific models and datasets. The authors propose a roadmap that includes native MEG pretraining, adaptation of EEG models, transfer from generic time‑series models, and multimodal integration with other neuroimaging and behavioral data, while emphasizing the need for coordinated infrastructure, rigorous evaluation, and responsible data‑sharing practices.

By Philipp Th\"olke, Hamza Abdelhedi, Yorguin Mantilla-Ramos, Fouad Lbakali, Oumayma Gharbi, Catherine Duclos, Annalisa Pascarella, Vanessa Hadid, Oiwi Parker Jones, Karim Jerbi
arXiv AI
Sep 12

RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection

RAMamba-Net is a new multimodal fusion network designed for auditory attention decoding (AAD) that combines EEG and electrooculography (EOG) signals. It uses a Mamba-enhanced band-aware convolutional Transformer to capture EEG band-specific patterns and long-range temporal dynamics, while a dual-branch encoder models EOG temporal and inter-channel dependencies. Cross‑modal attention and a reliability‑aware module estimate sample‑wise modality weights, improving fusion robustness and achieving a 5.76% accuracy gain over unimodal baselines on two AAD benchmarks.

By Xingyi He, Ziwei Wang, Dongrui Wu