arXiv:2609.35839v1 Announce Type: cross
Abstract: Handclaps provide an equipment-free excitation for room acoustics, but their unknown and variable source waveform makes room impulse response (RIR) e...
By Shih-Yu Lai, Kyung Yun Lee, Nils Meyer-Kahlen, Eloi Moliner, Bing-Yu Chen, Vesa V\"alim\"aki
arXiv:2609.38897v1 Announce Type: cross
Abstract: Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emp...
By Shivam Saini, Eric Bezzam, Georg G\"otz, Alessia Milo, Steinar Gu{\dh}j\'onsson, Konstantinos Gkanos, Finnur Pind, Daniel Gert Nielsen
The study examines how the realism of synthetic room impulse response (RIR) datasets influences the training of DeepFilterNet3 for single‑channel speech enhancement. By comparing a DNS4 image‑source‑method RIR set with a higher‑fidelity hybrid wave‑based and geometrical acoustics RIR set, the authors find that the more realistic dataset consistently improves objective speech enhancement metrics and significantly reduces ASR word error rates on unseen measured RIRs. The results suggest that overall realism in synthetic acoustic training data enhances DeepFilterNet3’s generalization to new environments.
By Alessia Milo, Georg G\"otz, Steinar Gu{\dh}j\'onsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind
arXiv:2607. 23293v1 Announce Type: cross Abstract: Image-source-method (ISM)-based room impulse response (RIR) simulation is a useful and physically interpretable tool for acoustic scene modeling, but full-order ISM becomes computationally expensive as the reflection order and room complexity increase.
By Shaoheng Xu, Chunyi Sun, Jihui Zhang, Amy Bastine, Prasanga N. Samarasinghe, Thushara D. Abhayapala
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordi...
The paper introduces asymmetric classifier‑free guidance (CFG) for target‑speaker ASR using Whisper, where a speaker‑conditioned branch predicts the target transcript and a speaker‑unconditioned branch predicts serialized multi‑speaker transcripts. CFG modulates the influence of speaker conditioning during decoding via a single guidance scale, which is first set globally on development data and then refined per utterance by a lightweight encoder‑based predictor while keeping the recognition model fixed. The resulting system yields up to 21.8% relative WER reduction over a condition‑only baseline and 5.6% over standard conditional decoding under domain shifts.
By Yiwen Guan, Jacob Whitehill