arXiv:2609.23951v1 Announce Type: new
Abstract: Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior w...
By Devangi Sharma, Sophia Judicke, Glenda Tan, Conrad Schaumburg, Shinji Watanabe
arXiv:2607. 06929v1 Announce Type: cross Abstract: Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual judgments.
By Sirui Zhang, Tianle Wang, Xinyi Tong, Peiyang Yu, Jishang Chen, Liangke Zhao, Haoxin Zhang, Duo Xu, Xin Jin, Feng Yu, Songchun Zhu
arXiv:2606. 10010v1 Announce Type: cross Abstract: Evaluating text-to-music (TTM) systems remains expensive because music impression (MI) and text alignment (TA) scores rely on human mean opinion scores (MOS).
By Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen
Chordonomicon is a new dataset of over 666,000 song-level symbolic chord progressions, each annotated with structural parts such as verse, chorus, and bridge, as well as genre and release date. The dataset was compiled by scraping user-generated progressions from multiple sources and shows strong similarity to established prior datasets. The authors also provide a reproducible benchmark suite for next chord prediction, evaluating RNN, GRU, and LSTM models across various context windows and data scales, and find that structural part annotations consistently improve prediction performance.
By Spyridon Kantarelis, Ioannis Liolitsas, Konstantinos Thomas, Vassilis Lyberatos, Edmund Dervakos, Giorgos Stamou
The paper evaluates five large language models as zero‑shot annotators of four social constructs—self‑esteem, self‑control, seeking belonging, and seeking recognition—in English song lyrics. It examines repeated‑measurement reliability, cross‑model convergence, and the transferability of consensus labels to supervised classification. Results show varying reliability across constructs, with self‑esteem being most stable and seeking recognition least stable, and indicate that consensus labels contain learnable signal for downstream tasks.
By E. Cho Smith, Samuel Ho, Dawn Laux
arXiv:2603. 00610v3 Announce Type: replace-cross Abstract: While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind.
By Yinghao Ma, Haiwen Xia, Hewei Gao, Weixiong Chen, Yuxin Ye, Yuchen Yang, Sungkyun Chang, Mingshuo Ding, Yizhi Li, Ruibin Yuan, Simon Dixon, Emmanouil Benetos
arXiv:2403.17612v3 Announce Type: replace
Abstract: Labeling corpora constitutes a bottleneck to create models for new tasks or domains. Large language models mitigate the issue with automatic corpus...
By Christopher Bagdon, Prathamesh Karmalker, Harsha Gurulingappa, Roman Klinger
The paper introduces a unified multitask learning framework for Music Emotion Recognition that simultaneously handles categorical and dimensional emotion labels, enabling training across multiple datasets. It leverages musical features such as key and chords, MERT embeddings, and employs knowledge distillation from teacher models to a student model to improve generalization. Experiments on MTG‑Jamendo, DEAM, PMEmo, and EmoMusic show that this approach outperforms state‑of‑the‑art models, including the top MediaEval 2021 entry.
By Jaeyong Kang, Dorien Herremans
arXiv:2606. 02638v1 Announce Type: cross Abstract: Recent advances in neural song generation have enabled high-quality synthesis from lyrics and global textual prompts.
By Yuejiao Wang, Zihao Ji, Pengfei Cai, Xu Li, Haorui Zheng, Zewen Song, Zhongliang Liu, Chen Zhang, Pengfei Wan
Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical poetry. However, domain-specific research on precise translation and affective-semantic understanding of classical poetry remains limited.
arXiv:2505. 18614v5 Announce Type: replace-cross Abstract: Lyrics translation requires both accurate semantic transfer and preservation of musical rhythm, syllabic structure, and poetic style.
By Woohyun Cho, Youngmin Kim, Sunghyun Lee, Youngjae Yu
arXiv:2606. 12392v1 Announce Type: cross Abstract: Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical poetry.
By Haotao Xie