arXiv AI

BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

arXiv:2604. 19532v3 Announce Type: replace-cross Abstract: Tokenizing music to fit the general framework of language models is a compelling challenge, especially considering the diverse symbolic structures in which music can be represented (e.

arXiv Machine Learning
Aug 31

How Far Should Tokenization Go? Predictive Effectiveness and Relational Losslessness

The paper proposes the Effectiveness–Losslessness Framework to guide tokenization in domains beyond language, using predictive codelength as a criterion. It introduces two boundaries: the Fact–Token Boundary, where observable structure should be encoded into tokens, and the Token–State Boundary, where context‑dependent relations should remain for model state rather than being pre‑tokenized. Experiments on symbolic music show that making musical time explicit and applying tonal‑frame canonicalization improve predictive performance, while fixed pitch coordinates and reversible BPE can increase predictive code length, indicating that carrier compaction alone does not guarantee better predictions.

By Yi Wang
arXiv Machine Learning
Sep 14

Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model

The paper introduces Musical Attention, a Transformer-based music generation model that incorporates meta-information such as bar numbers, key, signatures, and tempos into its attention mechanism. By representing each note with five events (pitch, bar number, onset, duration, velocity) plus three metadata elements, the model captures correlations among eight features, improving musical coherence and reducing repetition. Experiments show that Musical Attention outperforms prior methods like Full Attention and Strided Attention in coherence, variation, and overall quality, producing more diverse and harmonically consistent melodies.

By Shinnosuke Takasuka, Hideo Mukai
arXiv AI
Sep 15

DuoTok: Source-Aware Dual-Track Music Tokenization for Vocal-Accompaniment Generation

DuoTok is a source‑aware dual‑track music tokenizer designed for vocal‑accompaniment generation. It first learns a semantic audio representation via self‑supervised pretraining, then refines source‑aware structure with feature‑replacement noise and multi‑task supervision (spectral reconstruction, source separation regularization, and an ASR head for lyric alignment). The encoder is frozen and hard‑routed codebooks for vocals and accompaniment are learned, while a diffusion decoder restores fine acoustic detail from the discrete tokens, achieving a favorable predictability‑fidelity trade‑off at ultra‑low bitrate across public benchmarks.

By Rui Lin, Zhiyue Wu, Jiahe Lei, Kangdi Wang, Weixiong Chen, Junyu Dai, Tao Jiang
arXiv AI
Aug 19

Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System

The paper explains why GPT‑style language models fail to transfer directly to symbolic music. It argues that success in language comes from tokenization that compresses data by creating a coordinate system where recurring patterns become predictable. For music, the authors propose that tokenization must build a predictively effective, relationally lossless coordinate system—defining Fact–Token and Token–State boundaries—to enable compression without sacrificing contextual freedom. Controlled experiments confirm that proper coordinate construction improves predictive compressibility, whereas mere sequence compaction does not.

By Yi Wang
arXiv AI
Aug 20

Whole-Piece Training for Symbolic Music Language Models via Full-Horizon Compressed Recurrence

The paper introduces Whole-Piece Training for Symbolic Music Language Models using Full-Horizon Compressed Recurrence (FHCR), which maintains the full temporal horizon of recurrent memory while compressing its key-value representation to fit GPU limits. An evaluation diagnostic, KV-Reset Context Utilization (KRCU), demonstrates that full-horizon models retain long-range context beyond local windows, whereas limiting recurrent memory weakens this dependence. FHCR thus preserves long-range context utilization while significantly reducing recurrent memory cost, enabling efficient whole-piece modeling.

By Yungang Yi, Weihua Li, Matthew Kuo, Catherine Shi, Quan Bai
arXiv AI
Sep 17

CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling

The paper introduces CPR, a piano rendering framework that combines continuous autoregressive modeling with local flow matching and full‑sequence refinement. It predicts continuous hidden states, generates 24 kHz acoustic latents, and upsamples to 48 kHz, while new techniques BREPA and MT‑RoPE enhance musical semantics and cross‑modal alignment.

By Chong Jing, Junan Zhang, Zhizheng Wu
arXiv Machine Learning
Jul 17

MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

arXiv:2607. 14537v1 Announce Type: cross Abstract: Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures.

By Scott H. Hawley