arXiv:2511. 05550v3 Announce Type: replace-cross Abstract: Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio.
By Daniel Chenyu Lin, Michael Freeman, John Thickstun
arXiv:2607. 00641v1 Announce Type: cross Abstract: Advances in generative AI are rapidly increasing the quality and commercial value of generated music, and this progress depends on large catalogs of creators' recordings.
By Luyang Zhang, Xirui Jiang, Junwei Deng, Beibei Li, Jiaqi W. Ma, Chris Donahue
The paper introduces the MATCHA dataset, comprising 1,105 perceptual assessments from 83 experts on attribute-based music matches across five musical attributes—melody, harmony, rhythm, voice, and timbre. A triplet-based forced-choice experiment with 300 cases, including plagiarism, cover songs, and AI-generated music, was used to gather these judgments. Results show measurable agreement among participants and partial alignment with computational similarity measures, highlighting the need for perceptually grounded evaluation in generative AI for music.
By Roser Batlle-Roca, Woosung Choi, Joan Serr\`a, Fabio Morreale, Wei-Hsiang Liao, Xavier Serra, Emilia G\'omez, Yuki Mitsufuji
arXiv:2606. 04928v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed across diverse applications, raising critical questions for governance, accountability, and data provenance.
By Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Kaan Bayraktar, Roger Wattenhofer
arXiv:2609.39552v1 Announce Type: cross
Abstract: Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these pheno...
By Arhan Vohra, Choenden Kyirong, Laura Ib\'a\~nez-Mart\'inez, Mart\'in Rocamora
Chordonomicon is a new dataset of over 666,000 song-level symbolic chord progressions, each annotated with structural parts such as verse, chorus, and bridge, as well as genre and release date. The dataset was compiled by scraping user-generated progressions from multiple sources and shows strong similarity to established prior datasets. The authors also provide a reproducible benchmark suite for next chord prediction, evaluating RNN, GRU, and LSTM models across various context windows and data scales, and find that structural part annotations consistently improve prediction performance.
By Spyridon Kantarelis, Ioannis Liolitsas, Konstantinos Thomas, Vassilis Lyberatos, Edmund Dervakos, Giorgos Stamou
arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.
By Tawsif Ahmed, Andrej Radonjic, Gollam Rabby
arXiv:2603. 00610v3 Announce Type: replace-cross Abstract: While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind.
By Yinghao Ma, Haiwen Xia, Hewei Gao, Weixiong Chen, Yuxin Ye, Yuchen Yang, Sungkyun Chang, Mingshuo Ding, Yizhi Li, Ruibin Yuan, Simon Dixon, Emmanouil Benetos
MUUNRiver-Bench is a diagnostic benchmark for music retrieval that uses natural‑language instructions to define relevance for reference‑audio queries. It contains 3,440 tracks across 13 genres and 116 sub‑genres and covers seven tasks such as similar‑music, style‑preserving lyric‑rewriting, cover, and segment retrieval. Experiments with six models in eight configurations show that acoustic encoders favor local identity while text‑aligned encoders favor semantic relations, and that instruction‑aware and audio‑text fusion systems do not consistently outperform their backbones.
By Zhancheng Guo, Congren Dai, Shangda Wu, Jianhuai Hu, Danni Zhao, Xiaobing Li, Maosong Sun
The paper introduces REASONS, a benchmark of 12,723 sentence-level citation instances across 12 arXiv subject categories, to evaluate scientific citation attribution under different evidence conditions. It proposes a dual-metric framework—Abstention Rate (AR) and Hallucination Rate (HR)—to balance reliability and responsiveness. Experiments with proprietary and open-source LLMs across various prompting and retrieval settings show that advanced Retrieval-Augmented Generation (RAG) reduces hallucinations but increases abstention, while adversarial metadata can push hallucination rates above 85%. Human evaluation confirms a high ratio of factual hallucinations to acceptable paraphrases, underscoring the need for systems that can appropriately abstain under uncertainty.
By Deepa Tilwani, Yash Saxena, Seyedali Mohammadi, Ankur Padia, Edward Raff, Amit Sheth, Srinivasan Parthasarathy, Manas Gaur
The paper evaluates whether music‑text models truly capture fine‑grained musical meaning by introducing attribute‑swap perturbations that exchange properties such as timbre or order between instruments in a caption. Four contrastive models and one large audio‑language model were tested to see if they would score higher on the original caption than on the perturbed one. The results show that none of the contrastive models reliably distinguish the captions, and the audio‑language model’s advantage stems mainly from language priors, indicating that CLAP scores behave like a bag‑of‑words and fail to reflect attribute bindings.
By Yuan-Chiao Cheng, Alexander Lerch
The paper introduces REASONS, a benchmark comprising 12,723 sentence-level citation instances across 12 arXiv subject categories, to evaluate scientific citation attribution by large language models. It proposes a dual-metric framework—Abstention Rate (AR) and Hallucination Rate (HR)—to assess the trade-off between reliability and responsiveness. Experiments on proprietary and open-source LLMs under various prompting and retrieval settings show that advanced Retrieval-Augmented Generation (RAG) reduces hallucinations but may increase abstention, while retrieval-augmented variants often maintain near-zero abstention. Human evaluation reveals a high ratio of factual hallucinations to acceptable paraphrases, underscoring the need for systems that can appropriately abstain under uncertainty.
By Deepa Tilwani, Yash Saxena, Seyedali Mohammadi, Ankur Padia, Edward Raff, Amit Sheth, Srinivasan Parthasarathy, Manas Gaur