This survey reviews recent advances in converting non‑invasive EEG signals into images, text, and audio using generative AI techniques such as GANs, VAEs, transformers, and diffusion models. It summarizes datasets, feature‑encoding methods, evaluation metrics, and key challenges, noting that EEG‑to‑image models mainly use encoder‑decoder architectures, EEG‑to‑text leverages transformer language models, and EEG‑to‑audio maps signals to mel‑spectrograms for vocoder synthesis. The paper highlights the limitations of small, heterogeneous datasets, poor cross‑subject generalization, and the lack of standardized benchmarks, while providing open‑source resources to support reproducible research.
By Shreya Shukla, Jose Torres, Akshaj Murhekar, Christina Liu, Abhijit Mishra, Jacek Gwizdka, Shounak Roychowdhury
AFA‑Net introduces a differential attention mechanism to improve Auditory Attention Detection (AAD) from EEG signals, explicitly targeting noisy data. The framework outperforms existing deep learning models, achieving 96.8% accuracy within a 2‑second decision window while using fewer parameters. It represents one of the first approaches to actively mitigate EEG noise in AAD tasks.
By Philip H. Lee, Shreeram Suresh Chandra, Karan Thakkar, John H. L. Hansen
arXiv:2606. 24164v1 Announce Type: cross Abstract: Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies.
By Wonchul Shin, Inyong Choi, Kyogu Lee
Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance can be driven by trial-specific EEG structure that acts as shortcuts for target selection, leading to poor generalization on unseen trials.
arXiv:2606. 24087v1 Announce Type: new Abstract: Reconstructing continuous speech from scalp electroencephalography (EEG) remains fundamentally challenging.
By Wenhao Gao, Yifan Wang, Yijia Ma, Carl Yang, Wen Li, Chenyu You
The paper introduces a Cross-Subject Perceived Speech Decoding (CPSD) framework that tackles the challenge of decoding perceived speech from non‑invasive brain recordings across different subjects. CPSD uses a two‑stage training process: first, contrastive learning pre‑trains a source model on multiple subjects to capture shared representations; second, personal specialization fine‑tunes the model for a target subject by extracting consistent components and further training on that subject’s data. A Positional Encoding‑based Spatial Attention (PESA) module is added to remap MEG/EEG data into a standardized reference space, improving cross‑subject consistency. Evaluations on three datasets (Armeni 2022, PKUEEG 2025, Broderick 2018) show that CPSD outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top‑10 accuracy, demonstrating its effectiveness, efficiency, and robustness.
By Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen
EEGDM introduces a self‑supervised framework that uses latent diffusion models to generate EEG signals, moving beyond traditional masked reconstruction. The method employs an EEG encoder to produce a compact representation that conditions the diffusion denoising process, allowing joint optimization of encoder and generator. Experiments demonstrate that EEGDM can reconstruct high‑quality EEG, learn robust representations, and perform competitively on various downstream tasks.
By Shaocong Wang, Tong Liu, Yihan Li, Ming Li, Kairui Wen, Pei Yang, Wenqi Ji, Minjing Yu, Yong-Jin Liu
arXiv:2607. 05165v1 Announce Type: new Abstract: Non-invasive brain-to-speech decoding aims to restore communication to patients suffering from neurodegenerative disease, without the risks of neurosurgery.
By Benjamin Ballyk, Teyun Kwon, Miran \"Ozdogan, Oiwi Parker Jones
arXiv:2607. 25626v1 Announce Type: new Abstract: Decoding speech information directly from scalp electroencephalography (EEG) into text provides a potential non-invasive neural communication pathway for individuals with severe speech and motor impairments.
By Tian Zheng, Xurong Xie, Xinxin Zhu, Xiaolan Peng, Feng Tian
arXiv:2506.20354v3 Announce Type: replace-cross
Abstract: Learning from multi-variate time-series with heterogeneous channel configurations remains a fundamental challenge for deep neural networks, p...
By Francesco Carzaniga, Michael Hersche, Abu Sebastian, Kaspar Schindler, Abbas Rahimi
SHINE is a Sequential Hierarchical Integration Network designed to reconstruct speech envelope and Mel spectrogram from EEG and MEG recordings. It uses a residual sensor adapter, dilated-block states for temporal depth, and a target- and time-dependent gate to fuse hierarchical and attention-enhanced context predictions. Across two EEG and two MEG datasets, SHINE achieved the highest mean envelope and mean-Mel Pearson correlations among nine baseline methods and ranked second in the NeurIPS 2025 PNPL Competition’s speech-detection Extended Track.
By Xiran Xu, Yujie Yan, Songyi Li, Linze Zheng, Zifeng Zhang, Mochu Dong, Jing Chen
arXiv:2501. 09700v2 Announce Type: replace-cross Abstract: Electroencephalogram (EEG) signals have emerged as a promising modality for biometric identification.
By Ali Derakhshesh, Zahra Dehghanian, Reza Ebrahimpour, Hamid R. Rabiee