arXiv AI

MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering

arXiv:2608. 06638v1 Announce Type: cross Abstract: Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored.

arXiv Machine Learning
Jul 17

MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

arXiv:2607. 14537v1 Announce Type: cross Abstract: Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures.

By Scott H. Hawley
arXiv AI
Sep 7

Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes

The paper introduces a lightweight technique to steer the pitch content of audio generated by the Stable Audio Open diffusion model. A small convolutional probe (~125k parameters) is trained to decode frame‑level pitch‑class activations from the model’s latent space using paired audio and MIDI data. During inference, the frozen probe acts as a differentiable loss, guiding generation toward a user‑specified pitch‑class sequence without retraining the base model, and improves melodic coherence by 2.4× over the unguided baseline.

By Yushi Ye, Wilson Zheng, Yongyi Zang
arXiv AI
Jun 18

Closing the Loop: PID Feedback Control for Interpretable Activation Steering in Symbolic Music Generation

arXiv:2606. 18790v1 Announce Type: cross Abstract: Transformer-based architectures have significantly advanced the generation of complex symbolic sequences, yet a significant gap remains in achieving fine-grained, interpretable control over discrete signal attributes.

By Ioannis Prokopiou, Pantelis Vikatos, Maximos Kaliakatsos-Papakostas, Theodoros Giannakopoulos, Themos Stafylakis
arXiv Machine Learning
Aug 31

How Far Should Tokenization Go? Predictive Effectiveness and Relational Losslessness

The paper proposes the Effectiveness–Losslessness Framework to guide tokenization in domains beyond language, using predictive codelength as a criterion. It introduces two boundaries: the Fact–Token Boundary, where observable structure should be encoded into tokens, and the Token–State Boundary, where context‑dependent relations should remain for model state rather than being pre‑tokenized. Experiments on symbolic music show that making musical time explicit and applying tonal‑frame canonicalization improve predictive performance, while fixed pitch coordinates and reversible BPE can increase predictive code length, indicating that carrier compaction alone does not guarantee better predictions.

By Yi Wang
arXiv Machine Learning
Sep 23

Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning

SpanSynth-Edit is a flow‑matching model that enables MIDI‑guided synthesis and editing of multi‑instrument audio mixtures using low‑frame‑rate scalar‑quantised latents. It encodes instrument‑labelled note lifecycles as unordered event sets, pools them into a conditioning vector per audio‑latent frame, and uses contextual audio for instrument‑specific timbre guidance. The model supports editing by resynthesising target regions from revised MIDI and demonstrates competitive performance on single‑ and multi‑instrument benchmarks, including within‑frame onset control.

By Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos
arXiv Machine Learning
Sep 14

Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model

The paper introduces Musical Attention, a Transformer-based music generation model that incorporates meta-information such as bar numbers, key, signatures, and tempos into its attention mechanism. By representing each note with five events (pitch, bar number, onset, duration, velocity) plus three metadata elements, the model captures correlations among eight features, improving musical coherence and reducing repetition. Experiments show that Musical Attention outperforms prior methods like Full Attention and Strided Attention in coherence, variation, and overall quality, producing more diverse and harmonically consistent melodies.

By Shinnosuke Takasuka, Hideo Mukai
arXiv AI
Jun 26

Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training

arXiv:2512. 02652v2 Announce Type: replace-cross Abstract: Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance from a symbolic score, rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and language.

By Hong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia, Lan-Zhe Guo, Yu-Feng Li
arXiv AI
Aug 11

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation

arXiv:2608. 09035v1 Announce Type: cross Abstract: Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation.

By Shuyu Li, Kejun Zhang, Jiahe Lei, Shulei Ji, Zihao Wang, Jiaxing Yu, Wanying Wu, Lei Wang
arXiv AI
Sep 17

CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling

The paper introduces CPR, a piano rendering framework that combines continuous autoregressive modeling with local flow matching and full‑sequence refinement. It predicts continuous hidden states, generates 24 kHz acoustic latents, and upsamples to 48 kHz, while new techniques BREPA and MT‑RoPE enhance musical semantics and cross‑modal alignment.

By Chong Jing, Junan Zhang, Zhizheng Wu