arXiv Machine Learning

Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement

The study examines how the realism of synthetic room impulse response (RIR) datasets influences the training of DeepFilterNet3 for single‑channel speech enhancement. By comparing a DNS4 image‑source‑method RIR set with a higher‑fidelity hybrid wave‑based and geometrical acoustics RIR set, the authors find that the more realistic dataset consistently improves objective speech enhancement metrics and significantly reduces ASR word error rates on unseen measured RIRs. The results suggest that overall realism in synthetic acoustic training data enhances DeepFilterNet3’s generalization to new environments.

arXiv AI
Jun 30

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

arXiv:2606. 29031v1 Announce Type: cross Abstract: In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic speech recognition (ASR) without exposing sensitive customer recordings.

By Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso, S\'everin Baroudi, Shashi Kumar, Esa\'u Villatoro-Tello, Srikanth Madikeri, Manjunath K E, Old\v{r}ich Plchot, Kadri Hacio\u{g}lu, Petr Motlicek, Andreas Stolcke
arXiv AI
Sep 4

Test-time adaptation for speech enhancement with an autoregressive speech prior

The paper proposes a single‑utterance test‑time adaptation (TTA) method for speech enhancement that uses an autoregressive prior trained on clean speech latent representations from a neural audio codec. The adaptation regularizes a pretrained enhancement model by minimizing the Kullback‑Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments on multiple noisy speech datasets demonstrate consistent improvements in speech quality, especially when training and testing noise conditions differ.

By Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann
arXiv Machine Learning
Jul 28

PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response Simulation

arXiv:2607. 23293v1 Announce Type: cross Abstract: Image-source-method (ISM)-based room impulse response (RIR) simulation is a useful and physically interpretable tool for acoustic scene modeling, but full-order ISM becomes computationally expensive as the reflection order and room complexity increase.

By Shaoheng Xu, Chunyi Sun, Jihui Zhang, Amy Bastine, Prasanga N. Samarasinghe, Thushara D. Abhayapala
arXiv Machine Learning
Sep 11

Single Microphone Own Voice Detection based on Simulated Transfer Functions for Hearing Aids

The paper introduces a simulation-based method for detecting a user's own voice in hearing aids using only a single microphone. It employs a data augmentation strategy with simulated acoustic transfer functions to train a transformer classifier, achieving over 90% accuracy on both simulated and real-world recordings. The approach reduces hardware complexity and power consumption while maintaining robust performance across varied spatial conditions.

By Mathuranathan Mayuravaani, W. Bastiaan Kleijn, Andrew Lensen, Charlotte S{\o}rensen
arXiv AI
Jul 2

Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation

arXiv:2605. 07694v2 Announce Type: replace-cross Abstract: Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which components of the room impulse response (RIR) the model exploits and how performance depends on the recording conditions.

By Michael Neri, Archontis Politis, Tuomas Virtanen