Improving multichannel speech enhancement through accurate room-acoustic simulations
arXiv:2606. 31552v1 Announce Type: cross Abstract: Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement.
The study examines how the realism of synthetic room impulse response (RIR) datasets influences the training of DeepFilterNet3 for single‑channel speech enhancement. By comparing a DNS4 image‑source‑method RIR set with a higher‑fidelity hybrid wave‑based and geometrical acoustics RIR set, the authors find that the more realistic dataset consistently improves objective speech enhancement metrics and significantly reduces ASR word error rates on unseen measured RIRs. The results suggest that overall realism in synthetic acoustic training data enhances DeepFilterNet3’s generalization to new environments.
arXiv:2606. 31552v1 Announce Type: cross Abstract: Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement.
arXiv:2606. 29031v1 Announce Type: cross Abstract: In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic speech recognition (ASR) without exposing sensitive customer recordings.
arXiv:2604. 01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge.
arXiv:2509. 15210v2 Announce Type: replace-cross Abstract: Realistic sound simulation plays a critical role in many applications.
arXiv:2604. 14606v2 Announce Type: cross Abstract: Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates.
The paper proposes a single‑utterance test‑time adaptation (TTA) method for speech enhancement that uses an autoregressive prior trained on clean speech latent representations from a neural audio codec. The adaptation regularizes a pretrained enhancement model by minimizing the Kullback‑Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments on multiple noisy speech datasets demonstrate consistent improvements in speech quality, especially when training and testing noise conditions differ.
arXiv:2510. 20441v2 Announce Type: replace-cross Abstract: Neural audio codecs have largely promoted the application of language models (LMs) for speech applications.
arXiv:2608. 09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions.
arXiv:2607. 23293v1 Announce Type: cross Abstract: Image-source-method (ISM)-based room impulse response (RIR) simulation is a useful and physically interpretable tool for acoustic scene modeling, but full-order ISM becomes computationally expensive as the reflection order and room complexity increase.
The paper introduces a simulation-based method for detecting a user's own voice in hearing aids using only a single microphone. It employs a data augmentation strategy with simulated acoustic transfer functions to train a transformer classifier, achieving over 90% accuracy on both simulated and real-world recordings. The approach reduces hardware complexity and power consumption while maintaining robust performance across varied spatial conditions.
arXiv:2510. 16834v3 Announce Type: replace-cross Abstract: We present Schr\"odinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schr\"odinger Bridge (SB) training paradigm and the Mamba architecture.
arXiv:2605. 07694v2 Announce Type: replace-cross Abstract: Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which components of the room impulse response (RIR) the model exploits and how performance depends on the recording conditions.