arXiv:2607. 16736v1 Announce Type: cross Abstract: This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes.
By Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi, Tobias Morocutti, Gerhard Widmer
arXiv:2607. 01974v1 Announce Type: cross Abstract: This technical report describes our system for Task 1 of the DCASE 2026 Challenge, which aims to classify heterogeneous audio recordings according to the Broad Sound Taxonomy (BST).
By Beile Ning, Jiayi Yu, Zitong Wang, Yufei Hu, Wenjun Xu, Yuanhang Qian, Zhongxin Bai, Gongping Huang
arXiv:2607. 04383v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set.
By Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin
TUTTI is a new pre‑training framework for audio‑to‑score transcription that uses a large, fully synthetic multi‑instrument dataset generated by a symbolic music model. The approach trains a standard Transformer encoder‑decoder on these synthetic audio‑score pairs, producing a stronger foundational representation than single‑instrument training. When fine‑tuned on real datasets, TUTTI surpasses prior methods, achieving state‑of‑the‑art results and demonstrating strong cross‑instrument transferability.
By Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng, Chuanqi Yang, Xiaobing Li, Feng Yu, Maosong Sun
TEMPO is a unified model that adds temporally‑grounded capabilities to large audio‑language models, enabling timestamping of events, speakers, and sounds in audio, speech, and music. It introduces a supervised fine‑tuning stage featuring atomic timestamp tokens, a time‑aware projector with sinusoidal encodings, and a distance‑aware Gaussian loss, trained via a synthetic‑to‑real curriculum. Additionally, TEMPO employs reinforcement learning (GRPO) as a refinement step, and achieves state‑of‑the‑art performance on a benchmark of 10K samples across five timestamping tasks, surpassing Audio Flamingo Next and Qwen3‑Omni.
By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami, Dinesh Manocha
arXiv:2608. 06165v1 Announce Type: cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored.
By Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoro Mo, Yaolong Ju
The paper introduces Triage, a method that predicts the attention distribution of audio tokens before a language model processes them, enabling early pruning of less important tokens. By fitting a linear map to encoder outputs, Triage achieves high correlation (ρ ≥ 0.69) with full-model attention across eleven of thirteen large audio language models. Using this prediction, Triage compresses audio inputs while maintaining near‑full performance, outperforming baselines in transcription accuracy and significantly increasing the amount of audio that fits within a model’s context window.
By Kyoungjun Park, Yunzhe Li, Lili Qiu
Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. Relying solely on existing human-annotated datasets often...
arXiv:2607. 13571v1 Announce Type: cross Abstract: Sound event detection relies on frame-level strong labels whose annotation is expensive.
By Shiqi Zhang, Tuomas Virtanen
arXiv:2609.27389v1 Announce Type: cross
Abstract: Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is c...
By Yuxiang Wang, Shengbo Cai, Yingda Shen, Ming-Hao Hsu, Qinke Ni, Liqiang Zhang, Teddy Sun, Steve Yevs, Zhizheng Wu
arXiv:2502. 16584v2 Announce Type: replace-cross Abstract: Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs).
By Liumeng Xue, Ziya Zhou, Jiahao Pan, Zixuan Li, Shuai Fan, Yinghao Ma, Sitong Cheng, Dongchao Yang, Haohan Guo, Yujia Xiao, Xinsheng Wang, Zixuan Shen, Chuanbo Zhu, Xinshen Zhang, Tianchi Liu, Ruibin Yuan, Zeyue Tian, Haohe Liu, Xingjian Du, Emmanouil Benetos, Ge Zhang, Yike Guo, Wei Xue
The paper explores methods to mitigate catastrophic forgetting in incremental learning for sound event classification. It evaluates architectural and regularization strategies on FSD50K and AudioSet, finding that deeper layers, especially the classifier head, are most vulnerable. The most effective approach identified is fully freezing the feature extractor while fine‑tuning a dynamic head, which achieves minimal forgetting, stable training, and a balanced trade‑off between memory stability and learning plasticity.
By Riccardo Casciotti, Annamaria Mesaros