arXiv:2605. 03395v2 Announce Type: replace-cross Abstract: Music popularity prediction has attracted growing research interest, with relevance to artists, platforms, and recommendation systems.
By Jaavid Aktar Husain, Dorien Herremans
arXiv:2603. 00610v3 Announce Type: replace-cross Abstract: While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind.
By Yinghao Ma, Haiwen Xia, Hewei Gao, Weixiong Chen, Yuxin Ye, Yuchen Yang, Sungkyun Chang, Mingshuo Ding, Yizhi Li, Ruibin Yuan, Simon Dixon, Emmanouil Benetos
arXiv:2606. 06615v1 Announce Type: cross Abstract: Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries.
By Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha, Ramani Duraiswami
arXiv:2511. 05550v3 Announce Type: replace-cross Abstract: Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio.
By Daniel Chenyu Lin, Michael Freeman, John Thickstun
arXiv:2606. 00125v1 Announce Type: cross Abstract: Music recommendation systems typically treat songs as opaque tokens, relying on collaborative interaction histories which overlooks semantic or acoustic content.
By Srikar Prabhas Kandagatla, Sreehitha R. Narayana, Chandana Magapu, Swetha Mohan, Shamanth Kuthpadi, Hongjie Chen, Ryan A. Rossi, Franck Dernoncourt, Nesreen Ahmed
Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performances.
arXiv:2607. 27909v1 Announce Type: cross Abstract: Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes.
By Dmitrii Gavrilev, Ilya Borovik, Vladimir Viro
arXiv:2511. 23304v2 Announce Type: replace Abstract: In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE).
By Zijian Fu, Changsheng Lv, Xianlin Zhang, Mengshi Qi, Huadong Ma
Chordonomicon is a new dataset of over 666,000 song-level symbolic chord progressions, each annotated with structural parts such as verse, chorus, and bridge, as well as genre and release date. The dataset was compiled by scraping user-generated progressions from multiple sources and shows strong similarity to established prior datasets. The authors also provide a reproducible benchmark suite for next chord prediction, evaluating RNN, GRU, and LSTM models across various context windows and data scales, and find that structural part annotations consistently improve prediction performance.
By Spyridon Kantarelis, Ioannis Liolitsas, Konstantinos Thomas, Vassilis Lyberatos, Edmund Dervakos, Giorgos Stamou
The paper introduces a self‑supervised framework that maps text, audio, image, and video into a shared 256‑dimensional embedding space and uses iterative clustering to uncover aesthetic structure. It examines how AI’s cluster assignments diverge from human affective labels on a weakly supervised multimodal dataset. The study highlights implications for cross‑modal similarity, media organization for Retrieval‑Augmented Generation, and automated data labeling.
By Corey D. C. Heath
The paper explores how AI can develop its own aesthetic categorization of art across text, audio, image, and video without explicit labels. Using a self‑supervised framework, the authors embed these modalities into a shared 256‑dimensional space and iteratively cluster the data to uncover aesthetic structure. They compare the AI’s cluster assignments with human affective labels, highlighting divergences and discussing implications for cross‑modal similarity, media organization, and automated labeling.
arXiv:2607. 20253v1 Announce Type: cross Abstract: In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes.
By Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu