arXiv:2607. 18345v1 Announce Type: cross Abstract: Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs).
By David Rannaleet, Victor Gunnarsson, Bo Bernhardsson, Martin A. Skoglund, Emina Alickovic
arXiv:2606. 14120v1 Announce Type: cross Abstract: Auditory attention decoding (AAD) aims to infer the attended speaker from neural responses in multi-speaker acoustic environments and is a key problem for neuro-steered hearing systems.
By Ziwei Wang, Xingyi He, Tianwang Jia, Hongbin Wang, Dongrui Wu
RAMamba-Net is a new multimodal fusion network designed for auditory attention decoding (AAD) that combines EEG and electrooculography (EOG) signals. It uses a Mamba-enhanced band-aware convolutional Transformer to capture EEG band-specific patterns and long-range temporal dynamics, while a dual-branch encoder models EOG temporal and inter-channel dependencies. Cross‑modal attention and a reliability‑aware module estimate sample‑wise modality weights, improving fusion robustness and achieving a 5.76% accuracy gain over unimodal baselines on two AAD benchmarks.
By Xingyi He, Ziwei Wang, Dongrui Wu
The paper introduces Adaptive Anisotropic Attention (AAA), a method that splits self‑attention into temporal and spatial paths for axis‑structured signals like EEG. A learned gate combines the two paths for each token, and the resulting AXON model outperforms dense attention baselines on six EEG tasks and shows transfer to audio spectrograms. The study demonstrates that aligning attention with the natural axes of structured data provides a beneficial inductive bias.
By Mahir Jain, Parshva Runwal, Aditya Ray Mishra, Arvasu Kulkarni, Sandeep Singh, Siddharth Panwar
arXiv:2503. 00340v2 Announce Type: cross Abstract: Lightweight models are essential for real-time speech enhancement applications.
By Xiaobin Rong, Leyan Yang, Dahan Wang, Yuxiang Hu, Changbao Zhu, Kai Chen, Jing Lu
The paper introduces Adaptive Anisotropic Attention (AAA), a method that splits self‑attention into temporal and spatial paths for axis‑structured signals like EEG. A learned gate combines the two paths for each token, and the resulting AXON model outperforms dense attention baselines on six EEG tasks and shows benefits in audio spectrogram experiments. The study demonstrates that aligning attention with natural signal axes provides a useful inductive bias.