MusTBench is a music‑expert‑validated benchmark that evaluates temporal grounding in Large Audio‑Language Models (LALMs) through five temporally grounded question‑answering tasks. The paper also introduces MusT, a four‑stage optimization recipe—music encoder adaptation, LLM adaptation, supervised fine‑tuning, and RL‑based optimization—to improve temporal grounding. Experiments show that current LALMs struggle with precise temporal grounding, while MusT yields significant improvements, highlighting temporal grounding as a key missing capability in these models.
By Daeyong Kwon, Qiyu Wu, Shinobu Kuriya, Junghyun Koo, Shuyang Cui, Zhi Zhong, Wei-Hsiang Liao, Hiromi Wakaki, Yuki Mitsufuji
arXiv:2609.23416v1 Announce Type: cross
Abstract: Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape di...
By Zeyu Yang, Xinyu Zhang, Zibo Bi, Pei Zhang, Xize Cheng, Jin Xu, Baosong Yang, Satoshi Nakamura
TTM-Bench is a framework designed to benchmark text-to-music systems by establishing a common protocol for reproducible evaluation. It measures performance along two axes: musical-content alignment—assessed through semantic, genre, and musical-descriptor agreement scores against a shared musical specification—and computational efficiency, which includes generation latency, real-time factor, resource usage for local models, and cost for hosted services. A preliminary case study using TTM-Bench shows that higher alignment does not necessarily mean lower computational demands, underscoring the need for distinct, interpretable metrics.
By Giorgia Adorni, Michela Papandrea, Battista Rimoldi, Tiziano Leidi
arXiv:2511. 05550v3 Announce Type: replace-cross Abstract: Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio.
By Daniel Chenyu Lin, Michael Freeman, John Thickstun
arXiv:2602. 14612v4 Announce Type: replace-cross Abstract: Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding.
By Kartik Hegde, Arvind Krishna Sridhar, Naveen Vakada, Yinyi Guo, Erik Visser
arXiv:2603. 09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored.
By Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo, Ping-Le Tsai, Yen-Ting Piao, Hung-Wei Chen, Ting-Lin Hsiao, Yun-Man Hsu, Ke-Han Lu, Hung-yi Lee