The paper proposes the Effectiveness–Losslessness Framework to guide tokenization in domains beyond language, using predictive codelength as a criterion. It introduces two boundaries: the Fact–Token Boundary, where observable structure should be encoded into tokens, and the Token–State Boundary, where context‑dependent relations should remain for model state rather than being pre‑tokenized. Experiments on symbolic music show that making musical time explicit and applying tonal‑frame canonicalization improve predictive performance, while fixed pitch coordinates and reversible BPE can increase predictive code length, indicating that carrier compaction alone does not guarantee better predictions.
By Yi Wang
arXiv:2607. 11124v1 Announce Type: cross Abstract: Music creation is fundamentally a process of revision.
By Haoyu Gu, Lekai Qian, Haowu Zhou, Qi Liu, Shuai Wang
The paper introduces Musical Attention, a Transformer-based music generation model that incorporates meta-information such as bar numbers, key, signatures, and tempos into its attention mechanism. By representing each note with five events (pitch, bar number, onset, duration, velocity) plus three metadata elements, the model captures correlations among eight features, improving musical coherence and reducing repetition. Experiments show that Musical Attention outperforms prior methods like Full Attention and Strided Attention in coherence, variation, and overall quality, producing more diverse and harmonically consistent melodies.
By Shinnosuke Takasuka, Hideo Mukai
DuoTok is a source‑aware dual‑track music tokenizer designed for vocal‑accompaniment generation. It first learns a semantic audio representation via self‑supervised pretraining, then refines source‑aware structure with feature‑replacement noise and multi‑task supervision (spectral reconstruction, source separation regularization, and an ASR head for lyric alignment). The encoder is frozen and hard‑routed codebooks for vocals and accompaniment are learned, while a diffusion decoder restores fine acoustic detail from the discrete tokens, achieving a favorable predictability‑fidelity trade‑off at ultra‑low bitrate across public benchmarks.
By Rui Lin, Zhiyue Wu, Jiahe Lei, Kangdi Wang, Weixiong Chen, Junyu Dai, Tao Jiang
The paper explains why GPT‑style language models fail to transfer directly to symbolic music. It argues that success in language comes from tokenization that compresses data by creating a coordinate system where recurring patterns become predictable. For music, the authors propose that tokenization must build a predictively effective, relationally lossless coordinate system—defining Fact–Token and Token–State boundaries—to enable compression without sacrificing contextual freedom. Controlled experiments confirm that proper coordinate construction improves predictive compressibility, whereas mere sequence compaction does not.
By Yi Wang
The paper introduces Whole-Piece Training for Symbolic Music Language Models using Full-Horizon Compressed Recurrence (FHCR), which maintains the full temporal horizon of recurrent memory while compressing its key-value representation to fit GPU limits. An evaluation diagnostic, KV-Reset Context Utilization (KRCU), demonstrates that full-horizon models retain long-range context beyond local windows, whereas limiting recurrent memory weakens this dependence. FHCR thus preserves long-range context utilization while significantly reducing recurrent memory cost, enabling efficient whole-piece modeling.
By Yungang Yi, Weihua Li, Matthew Kuo, Catherine Shi, Quan Bai