arXiv:2608. 01481v1 Announce Type: new Abstract: Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.
By Ilia Semenkov, Daria Kleeva, Ivan Dakhtin, Zarina Maksudova, Alex Ossadtchi
The article outlines the emerging field of Magnetoencephalography (MEG) foundation models, explaining how these reusable, pretrained models can surpass traditional task‑specific decoding pipelines. It reviews current design choices—such as tokenization, sensor versus source representations, and self‑supervised objectives—and notes the limited number of existing MEG‑specific models and datasets. The authors propose a roadmap that includes native MEG pretraining, adaptation of EEG models, transfer from generic time‑series models, and multimodal integration with other neuroimaging and behavioral data, while emphasizing the need for coordinated infrastructure, rigorous evaluation, and responsible data‑sharing practices.
By Philipp Th\"olke, Hamza Abdelhedi, Yorguin Mantilla-Ramos, Fouad Lbakali, Oumayma Gharbi, Catherine Duclos, Annalisa Pascarella, Vanessa Hadid, Oiwi Parker Jones, Karim Jerbi
arXiv:2609.24095v1 Announce Type: new
Abstract: While deep learning has enabled language decoding from intracranial brain recordings, extending this capability to non-invasive recordings remains an u...
By Yueyang Li, Shuran Chen, Wai Ting Siok, Nizhuan Wang
While deep learning has enabled language decoding from intracranial brain recordings, extending this capability to non-invasive recordings remains an unresolved challenge. Decoding individual words fr...
The paper introduces a Cross-Subject Perceived Speech Decoding (CPSD) framework that tackles the challenge of decoding perceived speech from non‑invasive brain recordings across different subjects. CPSD uses a two‑stage training process: first, contrastive learning pre‑trains a source model on multiple subjects to capture shared representations; second, personal specialization fine‑tunes the model for a target subject by extracting consistent components and further training on that subject’s data. A Positional Encoding‑based Spatial Attention (PESA) module is added to remap MEG/EEG data into a standardized reference space, improving cross‑subject consistency. Evaluations on three datasets (Armeni 2022, PKUEEG 2025, Broderick 2018) show that CPSD outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top‑10 accuracy, demonstrating its effectiveness, efficiency, and robustness.
By Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen
The paper introduces the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which fuses fMRI and MEG data to decode perceived speech from non‑invasive brain signals. Comprehensive experiments show that SICMD improves Top‑1, Top‑10, and Rankacc scores by over 10%, 10%, and 1.7% respectively, while cutting training costs by 88.8% and 60.5% compared to existing multi‑subject and intra‑subject approaches. Visualizations further confirm the method’s effectiveness.
By Aoke Zhang, Jing Chen
arXiv:2609.40359v1 Announce Type: new
Abstract: We find that major reported improvements in decoding words from non-invasive brain recordings are largely reproducible without any brain data. In the i...
By Dulhan Jayalath, Oiwi Parker Jones
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
By Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin
arXiv:2607. 11801v1 Announce Type: cross Abstract: Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content.
By Yu-Han Huang, Chih-Kai Yang, Ke-Han Lu, An-Yu Cheng, Hung-yi Lee
LibriBrain100 is a new large‑scale MEG dataset for speech decoding that contains over 100 hours of high‑quality recordings while subjects listened to naturalistic continuous speech. The dataset more than doubles the size of the original LibriBrain release, with a record 80 hours from a single subject and additional 40‑minute recordings from 32 subjects. The authors demonstrate the value of deep within‑subject data and broad multi‑subject data by achieving state‑of‑the‑art word‑classification performance and showing that supervised fine‑tuning can compensate for limited per‑subject data, all supported by open‑source tools and a public competition leaderboard.
By Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran \"Ozdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones
arXiv:2609.05871v1 Announce Type: cross
Abstract: Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-s...
By Song-ha Jo, Sehyun Lee, Soyoon Kim, Jaesik Choi, Sanghyuk Choi
arXiv:2509. 24039v2 Announce Type: replace-cross Abstract: If topography is a fundamental feature of the brain, it should influence both how neurons are arranged in space (i.
By Haider Al-Tahan, Mayukh Deb, Jenelle Feather, N. Apurva Ratan Murty