arXiv:2606. 05852v1 Announce Type: cross Abstract: Text-to-speech (TTS) and singing voice synthesis (SVS) both aim to generate human vocal audio from symbolic inputs, but they impose different requirements on the generation process.
By Junjie Zheng, Huixin Xue, Shihong Ren, Chaofan Ding, Hao Liu, Zihao Chen
arXiv:2608. 13613v1 Announce Type: cross Abstract: Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions.
By Jiarui Hai, Karan Thakkar, Ke Chen, Yunyun Wang, Jiaqi Su, Rithesh Kumar, Mounya Elhilali, Zeyu Jin
arXiv:2608. 11590v1 Announce Type: cross Abstract: Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing.
By Haowei Lou, Hye-Young Paik, Dai Jia, Kai Li, Lina Yao
arXiv:2606. 24307v1 Announce Type: cross Abstract: Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from this domain due to its prohibitive inference latency and offline rendering paradigm.
By Baisen Wang, Chenxi Bao, Qisong Han
arXiv:2606. 31259v1 Announce Type: cross Abstract: Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step denoising.
By Binh Mai, Tran Quoc Bao Le, Hung Dinh, Cong Tran
arXiv:2606. 02638v1 Announce Type: cross Abstract: Recent advances in neural song generation have enabled high-quality synthesis from lyrics and global textual prompts.
By Yuejiao Wang, Zihao Ji, Pengfei Cai, Xu Li, Haorui Zheng, Zewen Song, Zhongliang Liu, Chen Zhang, Pengfei Wan
arXiv:2606. 03803v1 Announce Type: cross Abstract: We present LiveBand, a real-time system that generates high-fidelity music accompaniments to live audio input, respecting strict causal constraints.
By Marco Pasini, Javier Nistal, Mathias Rose Bjare, Stefan Lattner, George Fazekas
arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.
By Tawsif Ahmed, Andrej Radonjic, Gollam Rabby
arXiv:2606. 07293v1 Announce Type: cross Abstract: Speech Emotion Conversion (SEC) aims to transform the emotion of a source utterance into a target emotion while preserving content and speaker identity.
By Constantin Alexander Auga
arXiv:2606. 01703v1 Announce Type: cross Abstract: We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions.
By Jiashuo Yu, Yao Yao, Boyu Chen, Alex Wang
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.
In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content.