We study generative modeling of Bach-style symbolic piano music using a shared MIDI corpus and three model families: autoregressive LSTMs with attention, latent-variable models including recurrent VAEs and vector-quantized VAEs, and generative adversarial networks. We compare their ability to model polyphonic note sequences, learn useful latent representations, and generate stylistically coherent compositions.
arXiv:2608. 03050v1 Announce Type: cross Abstract: What is music style?
By Jingwei Zhao, Gus Xia, Ziyu Wang, Ye Wang
The paper introduces Musical Attention, a Transformer-based music generation model that incorporates meta-information such as bar numbers, key, signatures, and tempos into its attention mechanism. By representing each note with five events (pitch, bar number, onset, duration, velocity) plus three metadata elements, the model captures correlations among eight features, improving musical coherence and reducing repetition. Experiments show that Musical Attention outperforms prior methods like Full Attention and Strided Attention in coherence, variation, and overall quality, producing more diverse and harmonically consistent melodies.
By Shinnosuke Takasuka, Hideo Mukai
arXiv:2607. 14537v1 Announce Type: cross Abstract: Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures.
By Scott H. Hawley
The paper introduces CPR, a piano rendering framework that combines continuous autoregressive modeling with local flow matching and full‑sequence refinement. It predicts continuous hidden states, generates 24 kHz acoustic latents, and upsamples to 48 kHz, while new techniques BREPA and MT‑RoPE enhance musical semantics and cross‑modal alignment.
By Chong Jing, Junan Zhang, Zhizheng Wu
arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.
By Tawsif Ahmed, Andrej Radonjic, Gollam Rabby