This survey reviews recent advances in converting non‑invasive EEG signals into images, text, and audio using generative AI techniques such as GANs, VAEs, transformers, and diffusion models. It summarizes datasets, feature‑encoding methods, evaluation metrics, and key challenges, noting that EEG‑to‑image models mainly use encoder‑decoder architectures, EEG‑to‑text leverages transformer language models, and EEG‑to‑audio maps signals to mel‑spectrograms for vocoder synthesis. The paper highlights the limitations of small, heterogeneous datasets, poor cross‑subject generalization, and the lack of standardized benchmarks, while providing open‑source resources to support reproducible research.
By Shreya Shukla, Jose Torres, Akshaj Murhekar, Christina Liu, Abhijit Mishra, Jacek Gwizdka, Shounak Roychowdhury
arXiv:2511. 11686v4 Announce Type: replace Abstract: Speech enhancement (SE) requires high-fidelity reconstruction of clean speech that preserves linguistic and paralinguistic cues while maintaining high perceptual quality.
By Qing Yao, Lijian Gao, Qirong Mao, Ming Dong
arXiv:2607. 18345v1 Announce Type: cross Abstract: Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs).
By David Rannaleet, Victor Gunnarsson, Bo Bernhardsson, Martin A. Skoglund, Emina Alickovic
arXiv:2608. 00048v1 Announce Type: cross Abstract: Electroencephalography (EEG) generation is essential for alleviating data scarcity and enabling large scale neural modeling in brain computer interface applications.
By Boheng Liu, Ziyu Li, Chenghua Duan, Qing Li, Xia Wu
arXiv:2601. 09239v5 Announce Type: replace-cross Abstract: Speech tokenizers are a key building block of fully discrete Speech LLMs.
By Hanlin Zhang, Daxin Tan, Dehua Tao, Xiao Chen, Haochen Tan, Yunhe Li, Yuchen Cao, Linqi Song
arXiv:2606. 24164v1 Announce Type: cross Abstract: Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies.
By Wonchul Shin, Inyong Choi, Kyogu Lee
EEGDM introduces a self‑supervised framework that uses latent diffusion models to generate EEG signals, moving beyond traditional masked reconstruction. The method employs an EEG encoder to produce a compact representation that conditions the diffusion denoising process, allowing joint optimization of encoder and generator. Experiments demonstrate that EEGDM can reconstruct high‑quality EEG, learn robust representations, and perform competitively on various downstream tasks.
By Shaocong Wang, Tong Liu, Yihan Li, Ming Li, Kairui Wen, Pei Yang, Wenqi Ji, Minjing Yu, Yong-Jin Liu
arXiv:2607. 09134v1 Announce Type: cross Abstract: Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity.
By Sang-Hoon Lee, Ha-Yeong Choi
The paper introduces Corrective Forcing (CoF), a post‑training method that aligns diffusion and flow generative models for speech enhancement by training them on self‑generated rollout states. CoF corrects predictions toward ground truth under dynamic sampling schedules and regularizes local evolution with counterfactual transitions, applying a unified objective across both model types. Experiments on SB‑VE and OT‑CFM show improved perceptual quality, reconstruction fidelity, and robustness to varying sampling steps.
By Qing Yao, Lijian Gao, Qirong Mao
arXiv:2607. 29363v1 Announce Type: cross Abstract: Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation.
By Yi Luo, Rongzhi Gu, Jixun Yao
Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance can be driven by trial-specific EEG structure that acts as shortcuts for target selection, leading to poor generalization on unseen trials.
SHINE is a Sequential Hierarchical Integration Network designed to reconstruct speech envelope and Mel spectrogram from EEG and MEG recordings. It uses a residual sensor adapter, dilated-block states for temporal depth, and a target- and time-dependent gate to fuse hierarchical and attention-enhanced context predictions. Across two EEG and two MEG datasets, SHINE achieved the highest mean envelope and mean-Mel Pearson correlations among nine baseline methods and ranked second in the NeurIPS 2025 PNPL Competition’s speech-detection Extended Track.
By Xiran Xu, Yujie Yan, Songyi Li, Linze Zheng, Zifeng Zhang, Mochu Dong, Jing Chen