arXiv:2506. 01503v2 Announce Type: replace Abstract: With the rise of large pre-trained foundation models for automatic speech recognition new challenges appear.
By Benedikt Hilmes, Nick Rossenbach, Ralf Schl\"uter
arXiv:2606. 07387v1 Announce Type: new Abstract: State-of-the-art text-to-music generation systems rely on massive proprietary datasets and industrial-scale compute, making it impossible to disentangle architectural contributions from resource advantages.
By Yun-Chen Cheng, Tzu-Hung Huang, Chih-Pin Tan
The paper introduces Reward‑Tilted On‑Policy Distillation (RT‑OPD), a method that enhances acoustic grounding in audio‑language models by using a frozen teacher to generate a reward based on the contrast between token predictions with and without audio. This reward reshapes the teacher distribution for reverse‑KL distillation, encouraging students to rely more on acoustic evidence. Experiments on two compact students across three benchmarks show that RT‑OPD consistently outperforms vanilla OPD, and a 3B model trained with RT‑OPD achieves 72.72% accuracy on the MMAU benchmark, surpassing other 3B models and rivaling larger 7B and 8B models.
By Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji
arXiv:2601. 19919v2 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is one of the most effective paradigms for compressing large-scale foundation models into deployable architectures.
By Junseok Lee, Nahun Kim, Sangyong Lee, Chang-Jae Chun
arXiv:2603. 00610v3 Announce Type: replace-cross Abstract: While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind.
By Yinghao Ma, Haiwen Xia, Hewei Gao, Weixiong Chen, Yuxin Ye, Yuchen Yang, Sungkyun Chang, Mingshuo Ding, Yizhi Li, Ruibin Yuan, Simon Dixon, Emmanouil Benetos
arXiv:2609.27389v1 Announce Type: cross
Abstract: Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is c...
By Yuxiang Wang, Shengbo Cai, Yingda Shen, Ming-Hao Hsu, Qinke Ni, Liqiang Zhang, Teddy Sun, Steve Yevs, Zhizheng Wu