arXiv:2607. 16736v1 Announce Type: cross Abstract: This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes.
By Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi, Tobias Morocutti, Gerhard Widmer
arXiv:2608. 03742v1 Announce Type: cross Abstract: Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability.
By Sandy Abdo, Bill Kapralos, Priyamvada Tripathi, KC Collins, Adam Dubrowski
arXiv:2603. 02794v2 Announce Type: replace-cross Abstract: We present TVF (Time-Varying Filtering), an interpretable, low-latency speech enhancement model for real-time, on-device assistive hearing.
By Riccardo Rota, Kiril Ratmanski, Jozef Coldenhoff, Milos Cernak
arXiv:2606. 05121v1 Announce Type: cross Abstract: Audio is an inherently interactive modality, yet today's Large Audio Language Models (LALMs) are offline, and streaming audio models each handle only a single task such as streaming ASR or voice chatting.
By Zhifei Xie, Zihang Liu, Ze An, Xiaobin Hu, Yue Liao, Ziyang Ma, Dongchao Yang, Mingbao Lin, Deheng Ye, Shuicheng Yan, Chunyan Miao
arXiv:2606. 17775v1 Announce Type: cross Abstract: Efficient processing of continuous audio streams remains a key challenge for real-time and resource-constrained systems.
By Benjamin Hatton, Oliver Rhodes, Luca Peres
arXiv:2607. 09973v1 Announce Type: cross Abstract: Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows.
By M\'elodie Desbos, Yara Bahram, Eric Granger, Mohammadhadi Shateri
arXiv:2606. 24307v1 Announce Type: cross Abstract: Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from this domain due to its prohibitive inference latency and offline rendering paradigm.
By Baisen Wang, Chenxi Bao, Qisong Han
arXiv:2506. 20995v4 Announce Type: replace-cross Abstract: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis.
By Akio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji
arXiv:2510. 12947v3 Announce Type: replace-cross Abstract: Voice activity detection (VAD) serves as an early gate in voice-assistant pipelines for smart devices.
By Hamed Jafarzadeh Asl, Amin Edraki, Mahsa Ghazvini Nejad, Masoud Asgharian, Mohammadreza Sadeghi, Yuanhao Yu, Vahid Partovi Nia
arXiv:2503. 00340v2 Announce Type: cross Abstract: Lightweight models are essential for real-time speech enhancement applications.
By Xiaobin Rong, Leyan Yang, Dahan Wang, Yuxiang Hu, Changbao Zhu, Kai Chen, Jing Lu
arXiv:2606. 03803v1 Announce Type: cross Abstract: We present LiveBand, a real-time system that generates high-fidelity music accompaniments to live audio input, respecting strict causal constraints.
By Marco Pasini, Javier Nistal, Mathias Rose Bjare, Stefan Lattner, George Fazekas
arXiv:2608. 12615v1 Announce Type: cross Abstract: In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being.
By Cosmin Dragoiu, Nooshin Nabizadeh