arXiv:2607. 13587v1 Announce Type: cross Abstract: Automatic symbolic music analysis has made substantial progress, yet existing systems are typically designed for a single mode of use, such as full-score prediction, and therefore do not match the broader range of operations that arise in analysis workflows, including partial completion, local correction, and iterative refinement.
By Emmanouil Karystinaios, Johannes Hentschel, Markus Neuwirth, Gerhard Widmer
ONOTE is a unified framework that treats music as a scientifically structured domain of measurable cross-representation correspondences, focusing on omnimodal notation processing centered on sheet music. It introduces a test-only benchmark drawing from diverse musical sources—including staff, Jianpu, and tablature—across varied genres, instruments, and structural conditions, with aligned multimodal derivatives. The framework supports four complementary tasks—score understanding, notation conversion, audio transcription, and symbolic generation—while constructing a provenance-bearing proposition hypergraph from external music-theory materials for evidence retrieval and deterministic validity checks.
By Menghe Ma, Siqing Wei, Yuecheng Xing, Ziyue Zhu, Zhenghong Lin, Yaheng Wang, Fanhong Meng, Peijun Han, Luu Anh Tuan, Haoran Luo
arXiv:2511. 05550v3 Announce Type: replace-cross Abstract: Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio.
By Daniel Chenyu Lin, Michael Freeman, John Thickstun
arXiv:2606. 06615v1 Announce Type: cross Abstract: Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries.
By Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha, Ramani Duraiswami
arXiv:2607. 27909v1 Announce Type: cross Abstract: Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes.
By Dmitrii Gavrilev, Ilya Borovik, Vladimir Viro
MUUNRiver-Bench is a diagnostic benchmark for music retrieval that uses natural‑language instructions to define relevance for reference‑audio queries. It contains 3,440 tracks across 13 genres and 116 sub‑genres and covers seven tasks such as similar‑music, style‑preserving lyric‑rewriting, cover, and segment retrieval. Experiments with six models in eight configurations show that acoustic encoders favor local identity while text‑aligned encoders favor semantic relations, and that instruction‑aware and audio‑text fusion systems do not consistently outperform their backbones.
By Zhancheng Guo, Congren Dai, Shangda Wu, Jianhuai Hu, Danni Zhao, Xiaobing Li, Maosong Sun
Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performances.
arXiv:2607. 24873v1 Announce Type: new Abstract: Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts.
By Callie C. Liao, Duoduo Liao, Ellie L. Zhang
arXiv:2609.18585v1 Announce Type: cross
Abstract: Text-to-music (TTM) systems are increasingly used to generate musical audio from natural-language descriptions. Robust evaluation is therefore essent...
By Giorgia Adorni, Michela Papandrea, Battista Rimoldi, Tiziano Leidi
Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most existing systems follow a prompt-and-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas.
arXiv:2608.30940v1 Announce Type: cross
Abstract: Generative music systems are increasingly presented as tools that democratize music creation, yet their practical suitability for musicians remains u...
By Laura Ib\'a\~nez-Mart\'inez, Roser Batlle-Roca, Xavier Serra, Mart\'in Rocamora
MusTBench is a music‑expert‑validated benchmark that evaluates temporal grounding in Large Audio‑Language Models (LALMs) through five temporally grounded question‑answering tasks. The paper also introduces MusT, a four‑stage optimization recipe—music encoder adaptation, LLM adaptation, supervised fine‑tuning, and RL‑based optimization—to improve temporal grounding. Experiments show that current LALMs struggle with precise temporal grounding, while MusT yields significant improvements, highlighting temporal grounding as a key missing capability in these models.
By Daeyong Kwon, Qiyu Wu, Shinobu Kuriya, Junghyun Koo, Shuyang Cui, Zhi Zhong, Wei-Hsiang Liao, Hiromi Wakaki, Yuki Mitsufuji