arXiv:2608. 04378v1 Announce Type: cross Abstract: Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency.
By Scott H. Hawley
The paper introduces a lightweight technique to steer the pitch content of audio generated by the Stable Audio Open diffusion model. A small convolutional probe (~125k parameters) is trained to decode frame‑level pitch‑class activations from the model’s latent space using paired audio and MIDI data. During inference, the frozen probe acts as a differentiable loss, guiding generation toward a user‑specified pitch‑class sequence without retraining the base model, and improves melodic coherence by 2.4× over the unguided baseline.
By Yushi Ye, Wilson Zheng, Yongyi Zang
arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.
By Tawsif Ahmed, Andrej Radonjic, Gollam Rabby
arXiv:2608. 03920v1 Announce Type: cross Abstract: Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space.
By Zixun Guo, Simon Dixon
TrueMuse is a new benchmark designed to evaluate data attribution in text-to-music models. It consists of a controlled dataset created by fine‑tuning three diffusion‑based models on curated attribution samples, providing known attribution targets. The benchmark covers four settings—melodic structure, timbral characteristics, artist‑level style, and genre‑level patterns—across 133 attributes, 648 models, and 95,456 generated samples, and is used to assess existing black‑box attribution methods along several dimensions.
By Jiawei Yu, Jian Liu
DuoTok is a source‑aware dual‑track music tokenizer designed for vocal‑accompaniment generation. It first learns a semantic audio representation via self‑supervised pretraining, then refines source‑aware structure with feature‑replacement noise and multi‑task supervision (spectral reconstruction, source separation regularization, and an ASR head for lyric alignment). The encoder is frozen and hard‑routed codebooks for vocals and accompaniment are learned, while a diffusion decoder restores fine acoustic detail from the discrete tokens, achieving a favorable predictability‑fidelity trade‑off at ultra‑low bitrate across public benchmarks.
By Rui Lin, Zhiyue Wu, Jiahe Lei, Kangdi Wang, Weixiong Chen, Junyu Dai, Tao Jiang
arXiv:2607. 14537v1 Announce Type: cross Abstract: Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures.
By Scott H. Hawley
arXiv:2610.01864v1 Announce Type: cross
Abstract: How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Aut...
By Liwei Lin, Gus Xia
Project Qualia investigates whether experiential similarity between songs can be extracted from listening behavior. Using 1.29 billion scrobbles from 9,396 users, the authors trained a Word2Vec model (Song2Vec) on session data, then applied an artist‑residual procedure to isolate artist‑independent signals. The residual embeddings still contained strong cross‑artist similarity, forming coherent genre and era clusters, demonstrating that experiential structure exists beyond artist identity.
By Nizam Mohammed, Abu B. S. Rahman, Dimuthu D. K. Arachchige
arXiv:2608. 14819v1 Announce Type: cross Abstract: Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic.
By Angelos-Nikolaos Kanatas, Yuexuan Kong, Pablo Alonso-Jim\'enez, Xavier Serra, Dmitry Bogdanov
arXiv:2511. 05350v3 Announce Type: replace-cross Abstract: We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptually motivated losses, yields encodings that are structured according to a perceptual hierarchy.
By Mathias Rose Bjare, Giorgia Cantisani, Marco Pasini, Stefan Lattner, Gerhard Widmer
The paper introduces the MATCHA dataset, comprising 1,105 perceptual assessments from 83 experts on attribute-based music matches across five musical attributes—melody, harmony, rhythm, voice, and timbre. A triplet-based forced-choice experiment with 300 cases, including plagiarism, cover songs, and AI-generated music, was used to gather these judgments. Results show measurable agreement among participants and partial alignment with computational similarity measures, highlighting the need for perceptually grounded evaluation in generative AI for music.
By Roser Batlle-Roca, Woosung Choi, Joan Serr\`a, Fabio Morreale, Wei-Hsiang Liao, Xavier Serra, Emilia G\'omez, Yuki Mitsufuji