arXiv:2509. 15210v2 Announce Type: replace-cross Abstract: Realistic sound simulation plays a critical role in many applications.
By Chen Si, Qianyi Wu, Chaitanya Amballa, Romit Roy Choudhury
arXiv:2605. 07694v2 Announce Type: replace-cross Abstract: Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which components of the room impulse response (RIR) the model exploits and how performance depends on the recording conditions.
By Michael Neri, Archontis Politis, Tuomas Virtanen
arXiv:2606. 31552v1 Announce Type: cross Abstract: Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement.
By Georg G\"otz, Alessia Milo, Steinar Gu{\dh}j\'onsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind
arXiv:2606. 09677v1 Announce Type: cross Abstract: While discriminative models for multi-channel speech separation excel in reference-based metrics, they often exhibit suboptimal human listening quality.
By Dohwan Kim, Jung-Woo Choi
arXiv:2607. 04471v1 Announce Type: cross Abstract: Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization.
By Jakob Kienegger, Tal Peer, Sina Khanagha, Timo Gerkmann
arXiv:2607. 02119v1 Announce Type: cross Abstract: While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation.
By Haoran Wang, Jinchuan Tian, Siddhant Arora, Shinji Watanabe