arXiv AI By David Rannaleet, Victor Gunnarsson, Bo Bernhardsson, Martin A. Skoglund, Emina Alickovic

Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative Models

Read the original on arXiv AI →

arXiv:2607. 18345v1 Announce Type: cross Abstract: Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

A Survey on Bridging EEG Signals and Generative AI: From Image and Text to Beyond

This survey reviews recent advances in converting non‑invasive EEG signals into images, text, and audio using generative AI techniques such as GANs, VAEs, transformers, and diffusion models. It summarizes datasets, feature‑encoding methods, evaluation metrics, and key challenges, noting that EEG‑to‑image models mainly use encoder‑decoder architectures, EEG‑to‑text leverages transformer language models, and EEG‑to‑audio maps signals to mel‑spectrograms for vocoder synthesis. The paper highlights the limitations of small, heterogeneous datasets, poor cross‑subject generalization, and the lack of standardized benchmarks, while providing open‑source resources to support reproducible research.

By Shreya Shukla, Jose Torres, Akshaj Murhekar, Christina Liu, Abhijit Mishra, Jacek Gwizdka, Shounak Roychowdhury
arXiv Machine Learning
5d ago

AFA-Net: A Differential Attention Approach for Auditory Attention Detection

AFA‑Net introduces a differential attention mechanism to improve Auditory Attention Detection (AAD) from EEG signals, explicitly targeting noisy data. The framework outperforms existing deep learning models, achieving 96.8% accuracy within a 2‑second decision window while using fewer parameters. It represents one of the first approaches to actively mitigate EEG noise in AAD tasks.

By Philip H. Lee, Shreeram Suresh Chandra, Karan Thakkar, John H. L. Hansen
Hugging Face Trending Papers
Jun 23

Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training

Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance can be driven by trial-specific EEG structure that acts as shortcuts for target selection, leading to poor generalization on unseen trials.

arXiv AI
Aug 25

Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings

The paper introduces a Cross-Subject Perceived Speech Decoding (CPSD) framework that tackles the challenge of decoding perceived speech from non‑invasive brain recordings across different subjects. CPSD uses a two‑stage training process: first, contrastive learning pre‑trains a source model on multiple subjects to capture shared representations; second, personal specialization fine‑tunes the model for a target subject by extracting consistent components and further training on that subject’s data. A Positional Encoding‑based Spatial Attention (PESA) module is added to remap MEG/EEG data into a standardized reference space, improving cross‑subject consistency. Evaluations on three datasets (Armeni 2022, PKUEEG 2025, Broderick 2018) show that CPSD outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top‑10 accuracy, demonstrating its effectiveness, efficiency, and robustness.

By Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen