arXiv:2607. 14537v1 Announce Type: cross Abstract: Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures.
By Scott H. Hawley
arXiv:2607. 27909v1 Announce Type: cross Abstract: Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes.
By Dmitrii Gavrilev, Ilya Borovik, Vladimir Viro
arXiv:2608. 04378v1 Announce Type: cross Abstract: Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency.
By Scott H. Hawley
Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performances.
arXiv:2609.39552v1 Announce Type: cross
Abstract: Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these pheno...
By Arhan Vohra, Choenden Kyirong, Laura Ib\'a\~nez-Mart\'inez, Mart\'in Rocamora
arXiv:2608. 14819v1 Announce Type: cross Abstract: Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic.
By Angelos-Nikolaos Kanatas, Yuexuan Kong, Pablo Alonso-Jim\'enez, Xavier Serra, Dmitry Bogdanov
DuoTok is a source‑aware dual‑track music tokenizer designed for vocal‑accompaniment generation. It first learns a semantic audio representation via self‑supervised pretraining, then refines source‑aware structure with feature‑replacement noise and multi‑task supervision (spectral reconstruction, source separation regularization, and an ASR head for lyric alignment). The encoder is frozen and hard‑routed codebooks for vocals and accompaniment are learned, while a diffusion decoder restores fine acoustic detail from the discrete tokens, achieving a favorable predictability‑fidelity trade‑off at ultra‑low bitrate across public benchmarks.
By Rui Lin, Zhiyue Wu, Jiahe Lei, Kangdi Wang, Weixiong Chen, Junyu Dai, Tao Jiang
arXiv:2608.30974v1 Announce Type: cross
Abstract: Joint-Embedding Predictive Architecture (JEPA) has shown strong performance in learning rich representations through self-supervised prediction in la...
By Gabriel Meseguer-Brocal, Yuexuan Kong, Romain Hennequin
arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.
By Tawsif Ahmed, Andrej Radonjic, Gollam Rabby
The paper proposes the Effectiveness–Losslessness Framework to guide tokenization in domains beyond language, using predictive codelength as a criterion. It introduces two boundaries: the Fact–Token Boundary, where observable structure should be encoded into tokens, and the Token–State Boundary, where context‑dependent relations should remain for model state rather than being pre‑tokenized. Experiments on symbolic music show that making musical time explicit and applying tonal‑frame canonicalization improve predictive performance, while fixed pitch coordinates and reversible BPE can increase predictive code length, indicating that carrier compaction alone does not guarantee better predictions.
By Yi Wang
TUTTI is a new pre‑training framework for audio‑to‑score transcription that uses a large, fully synthetic multi‑instrument dataset generated by a symbolic music model. The approach trains a standard Transformer encoder‑decoder on these synthetic audio‑score pairs, producing a stronger foundational representation than single‑instrument training. When fine‑tuned on real datasets, TUTTI surpasses prior methods, achieving state‑of‑the‑art results and demonstrating strong cross‑instrument transferability.
By Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng, Chuanqi Yang, Xiaobing Li, Feng Yu, Maosong Sun
arXiv:2610.01864v1 Announce Type: cross
Abstract: How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Aut...
By Liwei Lin, Gus Xia