arXiv Machine Learning

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

arXiv:2608. 04378v1 Announce Type: cross Abstract: Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency.

arXiv Machine Learning
Jul 17

MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

arXiv:2607. 14537v1 Announce Type: cross Abstract: Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures.

By Scott H. Hawley
arXiv AI
Jun 26

Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training

arXiv:2512. 02652v2 Announce Type: replace-cross Abstract: Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance from a symbolic score, rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and language.

By Hong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia, Lan-Zhe Guo, Yu-Feng Li
arXiv AI
Sep 17

CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling

The paper introduces CPR, a piano rendering framework that combines continuous autoregressive modeling with local flow matching and full‑sequence refinement. It predicts continuous hidden states, generates 24 kHz acoustic latents, and upsamples to 48 kHz, while new techniques BREPA and MT‑RoPE enhance musical semantics and cross‑modal alignment.

By Chong Jing, Junan Zhang, Zhizheng Wu
Hugging Face Trending Papers
Jul 22

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content.

arXiv AI
Jul 23

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

arXiv:2607. 20253v1 Announce Type: cross Abstract: In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes.

By Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu
arXiv AI
Sep 15

DuoTok: Source-Aware Dual-Track Music Tokenization for Vocal-Accompaniment Generation

DuoTok is a source‑aware dual‑track music tokenizer designed for vocal‑accompaniment generation. It first learns a semantic audio representation via self‑supervised pretraining, then refines source‑aware structure with feature‑replacement noise and multi‑task supervision (spectral reconstruction, source separation regularization, and an ASR head for lyric alignment). The encoder is frozen and hard‑routed codebooks for vocals and accompaniment are learned, while a diffusion decoder restores fine acoustic detail from the discrete tokens, achieving a favorable predictability‑fidelity trade‑off at ultra‑low bitrate across public benchmarks.

By Rui Lin, Zhiyue Wu, Jiahe Lei, Kangdi Wang, Weixiong Chen, Junyu Dai, Tao Jiang
arXiv AI
Sep 7

Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes

The paper introduces a lightweight technique to steer the pitch content of audio generated by the Stable Audio Open diffusion model. A small convolutional probe (~125k parameters) is trained to decode frame‑level pitch‑class activations from the model’s latent space using paired audio and MIDI data. During inference, the frozen probe acts as a differentiable loss, guiding generation toward a user‑specified pitch‑class sequence without retraining the base model, and improves melodic coherence by 2.4× over the unguided baseline.

By Yushi Ye, Wilson Zheng, Yongyi Zang
arXiv AI
Jun 30

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

arXiv:2606. 30642v1 Announce Type: cross Abstract: Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts.

By Shun Lei, Huaicheng Zhang, Dapeng Wu, Yaoxun Xu, Lishi Zuo, Wei Tan, Hangting Chen, Guangzheng Li, Jianwei Yu, Zhiyong Wu, Dong Yu