We’re introducing Jukebox, a neural net that generates music, including rudimentary singing, as raw audio in a variety of genres and artist styles. We’re releasing the model weights and code, along with a tool to explore the generated samples.
arXiv:2607. 20253v1 Announce Type: cross Abstract: In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes.
By Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu
In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content.
arXiv:2502. 00023v2 Announce Type: replace-cross Abstract: Our research explores the development and application of musical agents, human-in-the-loop generative AI systems designed to support music performance and improvisation within co-creative spaces.
By Keon Ju M. Lee, Philippe Pasquier
arXiv:2608. 09035v1 Announce Type: cross Abstract: Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation.
By Shuyu Li, Kejun Zhang, Jiahe Lei, Shulei Ji, Zihao Wang, Jiaxing Yu, Wanying Wu, Lei Wang
The paper introduces Musical Attention, a Transformer-based music generation model that incorporates meta-information such as bar numbers, key, signatures, and tempos into its attention mechanism. By representing each note with five events (pitch, bar number, onset, duration, velocity) plus three metadata elements, the model captures correlations among eight features, improving musical coherence and reducing repetition. Experiments show that Musical Attention outperforms prior methods like Full Attention and Strided Attention in coherence, variation, and overall quality, producing more diverse and harmonically consistent melodies.
By Shinnosuke Takasuka, Hideo Mukai
arXiv:2607. 01849v1 Announce Type: cross Abstract: Musical performance involves executing a set of high-level musical instructions, yet recovering those instructions from the performance is a challenging inverse problem.
By Yewon Kim, Apurva Gandhi, David Chung, Graham Neubig, Chris Donahue
arXiv:2609.13291v1 Announce Type: cross
Abstract: We present ScorePrompts, an interactive system in which users upload a score, receive natural-language descriptions of its musical structure, ask que...
By Emmanouil Karystinaios, Gerhard Widmer
arXiv:2607. 05902v1 Announce Type: cross Abstract: Chamber music, as a highly precise multi-part interactive system, contains a logic of "role assignment and dynamic interaction" that provides an extremely valuable blueprint for exploring human-computer collaborative composition paradigms.
By Yakun Liu, Zhiyu Jin, Hai Luan, Dong Liu, Xiaonan Li
arXiv:2607. 11124v1 Announce Type: cross Abstract: Music creation is fundamentally a process of revision.
By Haoyu Gu, Lekai Qian, Haowu Zhou, Qi Liu, Shuai Wang
arXiv:2512. 02652v2 Announce Type: replace-cross Abstract: Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance from a symbolic score, rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and language.
By Hong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia, Lan-Zhe Guo, Yu-Feng Li
arXiv:2606. 12282v1 Announce Type: cross Abstract: Expressive performance rendering (EPR) aims to generate realistic performances constrained on sequences of notes.
By Dmitrii Gavrilev