arXiv:2512. 13998v3 Announce Type: replace-cross Abstract: Music Emotion Recognition (MER) is constrained by limited expert annotations and the need to establish robustness across heterogeneous corpora.
By Qilin Li, C. L. Philip Chen, Tong Zhang
arXiv:2603. 00610v3 Announce Type: replace-cross Abstract: While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind.
By Yinghao Ma, Haiwen Xia, Hewei Gao, Weixiong Chen, Yuxin Ye, Yuchen Yang, Sungkyun Chang, Mingshuo Ding, Yizhi Li, Ruibin Yuan, Simon Dixon, Emmanouil Benetos
The paper compares text‑based and feature‑based models for recognizing compound emotions in real‑world videos. It proposes textualizing non‑verbal cues from audio and visual modalities into text to leverage large language models, while feature‑based models directly combine extracted multimodal features. Experiments on the C‑EXPR‑DB dataset show that feature‑based models outperform textualization in the wild, though textual models can excel when rich transcripts are available.
By Nicolas Richet, Soufiane Belharbi, Haseeb Aslam, Meike Emilie Schadt, Manuela Gonz\'alez-Gonz\'alez, Gustave Cortal, Alessandro Lameiras Koerich, Marco Pedersoli, Alain Finkel, Simon Bacon, Eric Granger
arXiv:2607. 06929v1 Announce Type: cross Abstract: Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual judgments.
By Sirui Zhang, Tianle Wang, Xinyi Tong, Peiyang Yu, Jishang Chen, Liangke Zhao, Haoxin Zhang, Duo Xu, Xin Jin, Feng Yu, Songchun Zhu
arXiv:2608. 03920v1 Announce Type: cross Abstract: Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space.
By Zixun Guo, Simon Dixon
Chordonomicon is a new dataset of over 666,000 song-level symbolic chord progressions, each annotated with structural parts such as verse, chorus, and bridge, as well as genre and release date. The dataset was compiled by scraping user-generated progressions from multiple sources and shows strong similarity to established prior datasets. The authors also provide a reproducible benchmark suite for next chord prediction, evaluating RNN, GRU, and LSTM models across various context windows and data scales, and find that structural part annotations consistently improve prediction performance.
By Spyridon Kantarelis, Ioannis Liolitsas, Konstantinos Thomas, Vassilis Lyberatos, Edmund Dervakos, Giorgos Stamou