arXiv AI

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

arXiv:2606. 31338v1 Announce Type: cross Abstract: Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this reflects robust audio grounding or benchmark-specific shortcuts.

arXiv AI
Sep 2

MusTBench: Benchmarking and Advancing Temporal Grounding in Music LLMs

MusTBench is a music‑expert‑validated benchmark that evaluates temporal grounding in Large Audio‑Language Models (LALMs) through five temporally grounded question‑answering tasks. The paper also introduces MusT, a four‑stage optimization recipe—music encoder adaptation, LLM adaptation, supervised fine‑tuning, and RL‑based optimization—to improve temporal grounding. Experiments show that current LALMs struggle with precise temporal grounding, while MusT yields significant improvements, highlighting temporal grounding as a key missing capability in these models.

By Daeyong Kwon, Qiyu Wu, Shinobu Kuriya, Junghyun Koo, Shuyang Cui, Zhi Zhong, Wei-Hsiang Liao, Hiromi Wakaki, Yuki Mitsufuji
arXiv Machine Learning
Sep 17

TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking

TTM-Bench is a framework designed to benchmark text-to-music systems by establishing a common protocol for reproducible evaluation. It measures performance along two axes: musical-content alignment—assessed through semantic, genre, and musical-descriptor agreement scores against a shared musical specification—and computational efficiency, which includes generation latency, real-time factor, resource usage for local models, and cost for hosted services. A preliminary case study using TTM-Bench shows that higher alignment does not necessarily mean lower computational demands, underscoring the need for distinct, interpretable metrics.

By Giorgia Adorni, Michela Papandrea, Battista Rimoldi, Tiziano Leidi
arXiv AI
Sep 2

TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data

TUTTI is a new pre‑training framework for audio‑to‑score transcription that uses a large, fully synthetic multi‑instrument dataset generated by a symbolic music model. The approach trains a standard Transformer encoder‑decoder on these synthetic audio‑score pairs, producing a stronger foundational representation than single‑instrument training. When fine‑tuned on real datasets, TUTTI surpasses prior methods, achieving state‑of‑the‑art results and demonstrating strong cross‑instrument transferability.

By Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng, Chuanqi Yang, Xiaobing Li, Feng Yu, Maosong Sun
arXiv AI
Jul 7

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

arXiv:2607. 04383v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set.

By Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin