Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2. 0 audio embeddings.
arXiv:2609.24095v1 Announce Type: new
Abstract: While deep learning has enabled language decoding from intracranial brain recordings, extending this capability to non-invasive recordings remains an u...
By Yueyang Li, Shuran Chen, Wai Ting Siok, Nizhuan Wang
While deep learning has enabled language decoding from intracranial brain recordings, extending this capability to non-invasive recordings remains an unresolved challenge. Decoding individual words fr...
The paper introduces a Cross-Subject Perceived Speech Decoding (CPSD) framework that tackles the challenge of decoding perceived speech from non‑invasive brain recordings across different subjects. CPSD uses a two‑stage training process: first, contrastive learning pre‑trains a source model on multiple subjects to capture shared representations; second, personal specialization fine‑tunes the model for a target subject by extracting consistent components and further training on that subject’s data. A Positional Encoding‑based Spatial Attention (PESA) module is added to remap MEG/EEG data into a standardized reference space, improving cross‑subject consistency. Evaluations on three datasets (Armeni 2022, PKUEEG 2025, Broderick 2018) show that CPSD outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top‑10 accuracy, demonstrating its effectiveness, efficiency, and robustness.
By Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen
The article outlines the emerging field of Magnetoencephalography (MEG) foundation models, explaining how these reusable, pretrained models can surpass traditional task‑specific decoding pipelines. It reviews current design choices—such as tokenization, sensor versus source representations, and self‑supervised objectives—and notes the limited number of existing MEG‑specific models and datasets. The authors propose a roadmap that includes native MEG pretraining, adaptation of EEG models, transfer from generic time‑series models, and multimodal integration with other neuroimaging and behavioral data, while emphasizing the need for coordinated infrastructure, rigorous evaluation, and responsible data‑sharing practices.
By Philipp Th\"olke, Hamza Abdelhedi, Yorguin Mantilla-Ramos, Fouad Lbakali, Oumayma Gharbi, Catherine Duclos, Annalisa Pascarella, Vanessa Hadid, Oiwi Parker Jones, Karim Jerbi
The paper introduces the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which fuses fMRI and MEG data to decode perceived speech from non‑invasive brain signals. Comprehensive experiments show that SICMD improves Top‑1, Top‑10, and Rankacc scores by over 10%, 10%, and 1.7% respectively, while cutting training costs by 88.8% and 60.5% compared to existing multi‑subject and intra‑subject approaches. Visualizations further confirm the method’s effectiveness.
By Aoke Zhang, Jing Chen