The paper introduces a Cross-Subject Perceived Speech Decoding (CPSD) framework that tackles the challenge of decoding perceived speech from non‑invasive brain recordings across different subjects. CPSD uses a two‑stage training process: first, contrastive learning pre‑trains a source model on multiple subjects to capture shared representations; second, personal specialization fine‑tunes the model for a target subject by extracting consistent components and further training on that subject’s data. A Positional Encoding‑based Spatial Attention (PESA) module is added to remap MEG/EEG data into a standardized reference space, improving cross‑subject consistency. Evaluations on three datasets (Armeni 2022, PKUEEG 2025, Broderick 2018) show that CPSD outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top‑10 accuracy, demonstrating its effectiveness, efficiency, and robustness.
By Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen
arXiv:2506.20354v3 Announce Type: replace-cross
Abstract: Learning from multi-variate time-series with heterogeneous channel configurations remains a fundamental challenge for deep neural networks, p...
By Francesco Carzaniga, Michael Hersche, Abu Sebastian, Kaspar Schindler, Abbas Rahimi
arXiv:2606. 24087v1 Announce Type: new Abstract: Reconstructing continuous speech from scalp electroencephalography (EEG) remains fundamentally challenging.
By Wenhao Gao, Yifan Wang, Yijia Ma, Carl Yang, Wen Li, Chenyu You
arXiv:2510. 15371v2 Announce Type: replace-cross Abstract: Classification of electroencephalogram (EEG) signals obtained during motor imagery (MI) has substantial application potential, including communication assistance and rehabilitation support for patients with motor impairments.
By Shuntaro Suzuki, Shunya Nagashima, Komei Sugiura
arXiv:2607. 09543v1 Announce Type: new Abstract: Self-supervised pretrained foundation models (FM) have shown early promise for non-invasive electroencephalogram (EEG) decoding applications.
By Gabriel Mahuas, Victoria Shevchenko, Ugo Tanielian, Yassir Bendou, Richard Gao
The paper introduces the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which fuses fMRI and MEG data to decode perceived speech from non‑invasive brain signals. Comprehensive experiments show that SICMD improves Top‑1, Top‑10, and Rankacc scores by over 10%, 10%, and 1.7% respectively, while cutting training costs by 88.8% and 60.5% compared to existing multi‑subject and intra‑subject approaches. Visualizations further confirm the method’s effectiveness.
By Aoke Zhang, Jing Chen
arXiv:2607. 21402v1 Announce Type: new Abstract: Self-supervised foundation models have recently shown strong potential for electroencephalogram (EEG)-based analysis.
By Tao Zhou, Jing Han, Lingyu Shu, Zixing Zhang
arXiv:2606. 24164v1 Announce Type: cross Abstract: Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies.
By Wonchul Shin, Inyong Choi, Kyogu Lee
Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance can be driven by trial-specific EEG structure that acts as shortcuts for target selection, leading to poor generalization on unseen trials.
arXiv:2607. 18345v1 Announce Type: cross Abstract: Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs).
By David Rannaleet, Victor Gunnarsson, Bo Bernhardsson, Martin A. Skoglund, Emina Alickovic
LEAD is a gated temporal‑spatial Transformer foundation model designed for EEG‑based Alzheimer's disease detection. It was trained on the world’s largest EEG‑AD corpus of 2,238 subjects and uses a subject‑regularized strategy and medical contrastive learning across 13 datasets. LEAD outperforms existing methods on five downstream AD datasets, achieving the best average ranking across 20 evaluations.
By Yihe Wang, Nan Huang, Nadia Mammone, Marco Cecchi, Xiang Zhang
This survey reviews recent advances in converting non‑invasive EEG signals into images, text, and audio using generative AI techniques such as GANs, VAEs, transformers, and diffusion models. It summarizes datasets, feature‑encoding methods, evaluation metrics, and key challenges, noting that EEG‑to‑image models mainly use encoder‑decoder architectures, EEG‑to‑text leverages transformer language models, and EEG‑to‑audio maps signals to mel‑spectrograms for vocoder synthesis. The paper highlights the limitations of small, heterogeneous datasets, poor cross‑subject generalization, and the lack of standardized benchmarks, while providing open‑source resources to support reproducible research.
By Shreya Shukla, Jose Torres, Akshaj Murhekar, Christina Liu, Abhijit Mishra, Jacek Gwizdka, Shounak Roychowdhury