arXiv:2607. 04471v1 Announce Type: cross Abstract: Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization.
By Jakob Kienegger, Tal Peer, Sina Khanagha, Timo Gerkmann
arXiv:2606. 18664v1 Announce Type: cross Abstract: Reliable sound source localization is fundamental to robot audition, enabling autonomous robots to perceive spatial cues and operate effectively in dynamic environments.
By Yizhuo Yang, Junqiao Fan, Shenghai Yuan, Lihua Xie
arXiv:2606. 09677v1 Announce Type: cross Abstract: While discriminative models for multi-channel speech separation excel in reference-based metrics, they often exhibit suboptimal human listening quality.
By Dohwan Kim, Jung-Woo Choi
arXiv:2607. 01295v1 Announce Type: cross Abstract: Acoustic imaging visualization is a core methodology in acoustics, enabling spatial analysis of sound sources and acoustic scenes.
By Marianthi Adamopoulou, Parthasaarathy Sudarsanam, David Diaz-Guerra, Meng Jiang, Archontis Politis, Seyed Jalaleddin Mousavirad, Tuomas Virtanen, Jan Lundgren
The paper proposes a single‑utterance test‑time adaptation (TTA) method for speech enhancement that uses an autoregressive prior trained on clean speech latent representations from a neural audio codec. The adaptation regularizes a pretrained enhancement model by minimizing the Kullback‑Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments on multiple noisy speech datasets demonstrate consistent improvements in speech quality, especially when training and testing noise conditions differ.
By Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann
arXiv:2608. 15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output.
By Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov
Reliable sound source localization is fundamental to robot audition, enabling autonomous robots to perceive spatial cues and operate effectively in dynamic environments. Classical methods such as Multiple Signal Classification (MUSIC) offer strong theoretical foundations but degrade under low signal-to-noise ratios.
The paper introduces a simulation-based method for detecting a user's own voice in hearing aids using only a single microphone. It employs a data augmentation strategy with simulated acoustic transfer functions to train a transformer classifier, achieving over 90% accuracy on both simulated and real-world recordings. The approach reduces hardware complexity and power consumption while maintaining robust performance across varied spatial conditions.
By Mathuranathan Mayuravaani, W. Bastiaan Kleijn, Andrew Lensen, Charlotte S{\o}rensen
arXiv:2609.35839v1 Announce Type: cross
Abstract: Handclaps provide an equipment-free excitation for room acoustics, but their unknown and variable source waveform makes room impulse response (RIR) e...
By Shih-Yu Lai, Kyung Yun Lee, Nils Meyer-Kahlen, Eloi Moliner, Bing-Yu Chen, Vesa V\"alim\"aki
arXiv:2604. 24199v4 Announce Type: replace-cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem.
By Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson
The paper addresses challenges in extracting target and multiple speakers from real conversational speech, noting that real conversations contain more silence and enrolment samples that differ from the target speech. It introduces a new loss function that reduces the impact of excess silence during training, yielding improvements in STOI (from 0.55 to 0.60) and frequency‑weighted segmental SNR (from 4.35 to 5.12). The study also investigates how mismatches between enrolment and target speech affect performance.
By Robert Sutherland, Stefan Goetze, Jon Barker
arXiv:2606. 31552v1 Announce Type: cross Abstract: Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement.
By Georg G\"otz, Alessia Milo, Steinar Gu{\dh}j\'onsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind