arXiv Machine Learning

Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition

arXiv:2607. 22413v1 Announce Type: cross Abstract: Sample retrieval tools can help composers find harmonically compatible material, but querying from a fixed reference sample becomes less informative as arrangements evolve and the harmonic context shifts with each musical decision.

arXiv AI
Sep 17

CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling

The paper introduces CPR, a piano rendering framework that combines continuous autoregressive modeling with local flow matching and full‑sequence refinement. It predicts continuous hidden states, generates 24 kHz acoustic latents, and upsamples to 48 kHz, while new techniques BREPA and MT‑RoPE enhance musical semantics and cross‑modal alignment.

By Chong Jing, Junan Zhang, Zhizheng Wu
arXiv Machine Learning
Jul 17

MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

arXiv:2607. 14537v1 Announce Type: cross Abstract: Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures.

By Scott H. Hawley
arXiv AI
Sep 16

MUUNRiver-Bench: Diagnosing Relation-Dependent Music Retrieval with Multimodal Instructions

MUUNRiver-Bench is a diagnostic benchmark for music retrieval that uses natural‑language instructions to define relevance for reference‑audio queries. It contains 3,440 tracks across 13 genres and 116 sub‑genres and covers seven tasks such as similar‑music, style‑preserving lyric‑rewriting, cover, and segment retrieval. Experiments with six models in eight configurations show that acoustic encoders favor local identity while text‑aligned encoders favor semantic relations, and that instruction‑aware and audio‑text fusion systems do not consistently outperform their backbones.

By Zhancheng Guo, Congren Dai, Shangda Wu, Jianhuai Hu, Danni Zhao, Xiaobing Li, Maosong Sun
arXiv Machine Learning
Sep 11

Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data

Project Qualia investigates whether experiential similarity between songs can be extracted from listening behavior. Using 1.29 billion scrobbles from 9,396 users, the authors trained a Word2Vec model (Song2Vec) on session data, then applied an artist‑residual procedure to isolate artist‑independent signals. The residual embeddings still contained strong cross‑artist similarity, forming coherent genre and era clusters, demonstrating that experiential structure exists beyond artist identity.

By Nizam Mohammed, Abu B. S. Rahman, Dimuthu D. K. Arachchige
arXiv AI
Sep 1

TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

TEMPO is a unified model that adds temporally‑grounded capabilities to large audio‑language models, enabling timestamping of events, speakers, and sounds in audio, speech, and music. It introduces a supervised fine‑tuning stage featuring atomic timestamp tokens, a time‑aware projector with sinusoidal encodings, and a distance‑aware Gaussian loss, trained via a synthetic‑to‑real curriculum. Additionally, TEMPO employs reinforcement learning (GRPO) as a refinement step, and achieves state‑of‑the‑art performance on a benchmark of 10K samples across five timestamping tasks, surpassing Audio Flamingo Next and Qwen3‑Omni.

By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami, Dinesh Manocha
arXiv AI
Aug 25

ONOTE: Hypergraph-Grounded Omnimodal Reasoning for Computational Music Science

ONOTE is a unified framework that treats music as a scientifically structured domain of measurable cross-representation correspondences, focusing on omnimodal notation processing centered on sheet music. It introduces a test-only benchmark drawing from diverse musical sources—including staff, Jianpu, and tablature—across varied genres, instruments, and structural conditions, with aligned multimodal derivatives. The framework supports four complementary tasks—score understanding, notation conversion, audio transcription, and symbolic generation—while constructing a provenance-bearing proposition hypergraph from external music-theory materials for evidence retrieval and deterministic validity checks.

By Menghe Ma, Siqing Wei, Yuecheng Xing, Ziyue Zhu, Zhenghong Lin, Yaheng Wang, Fanhong Meng, Peijun Han, Luu Anh Tuan, Haoran Luo