arXiv AI By Yujun Lee, Joonhyeok Shin, Hyoeun Kim, Kyuhong Shim

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

Read the original on arXiv AI →

arXiv:2606. 31338v1 Announce Type: cross Abstract: Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this reflects robust audio grounding or benchmark-specific shortcuts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 2

MusTBench: Benchmarking and Advancing Temporal Grounding in Music LLMs

MusTBench is a music‑expert‑validated benchmark that evaluates temporal grounding in Large Audio‑Language Models (LALMs) through five temporally grounded question‑answering tasks. The paper also introduces MusT, a four‑stage optimization recipe—music encoder adaptation, LLM adaptation, supervised fine‑tuning, and RL‑based optimization—to improve temporal grounding. Experiments show that current LALMs struggle with precise temporal grounding, while MusT yields significant improvements, highlighting temporal grounding as a key missing capability in these models.

By Daeyong Kwon, Qiyu Wu, Shinobu Kuriya, Junghyun Koo, Shuyang Cui, Zhi Zhong, Wei-Hsiang Liao, Hiromi Wakaki, Yuki Mitsufuji
arXiv Machine Learning
Sep 17

TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking

TTM-Bench is a framework designed to benchmark text-to-music systems by establishing a common protocol for reproducible evaluation. It measures performance along two axes: musical-content alignment—assessed through semantic, genre, and musical-descriptor agreement scores against a shared musical specification—and computational efficiency, which includes generation latency, real-time factor, resource usage for local models, and cost for hosted services. A preliminary case study using TTM-Bench shows that higher alignment does not necessarily mean lower computational demands, underscoring the need for distinct, interpretable metrics.

By Giorgia Adorni, Michela Papandrea, Battista Rimoldi, Tiziano Leidi