arXiv:2511. 05550v3 Announce Type: replace-cross Abstract: Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio.
By Daniel Chenyu Lin, Michael Freeman, John Thickstun
arXiv:2607. 00641v1 Announce Type: cross Abstract: Advances in generative AI are rapidly increasing the quality and commercial value of generated music, and this progress depends on large catalogs of creators' recordings.
By Luyang Zhang, Xirui Jiang, Junwei Deng, Beibei Li, Jiaqi W. Ma, Chris Donahue
The paper introduces the MATCHA dataset, comprising 1,105 perceptual assessments from 83 experts on attribute-based music matches across five musical attributes—melody, harmony, rhythm, voice, and timbre. A triplet-based forced-choice experiment with 300 cases, including plagiarism, cover songs, and AI-generated music, was used to gather these judgments. Results show measurable agreement among participants and partial alignment with computational similarity measures, highlighting the need for perceptually grounded evaluation in generative AI for music.
By Roser Batlle-Roca, Woosung Choi, Joan Serr\`a, Fabio Morreale, Wei-Hsiang Liao, Xavier Serra, Emilia G\'omez, Yuki Mitsufuji
arXiv:2606. 04928v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed across diverse applications, raising critical questions for governance, accountability, and data provenance.
By Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Kaan Bayraktar, Roger Wattenhofer
arXiv:2609.39552v1 Announce Type: cross
Abstract: Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these pheno...
By Arhan Vohra, Choenden Kyirong, Laura Ib\'a\~nez-Mart\'inez, Mart\'in Rocamora
Chordonomicon is a new dataset of over 666,000 song-level symbolic chord progressions, each annotated with structural parts such as verse, chorus, and bridge, as well as genre and release date. The dataset was compiled by scraping user-generated progressions from multiple sources and shows strong similarity to established prior datasets. The authors also provide a reproducible benchmark suite for next chord prediction, evaluating RNN, GRU, and LSTM models across various context windows and data scales, and find that structural part annotations consistently improve prediction performance.
By Spyridon Kantarelis, Ioannis Liolitsas, Konstantinos Thomas, Vassilis Lyberatos, Edmund Dervakos, Giorgos Stamou