arXiv AI
Sep 7

Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes

The paper introduces a lightweight technique to steer the pitch content of audio generated by the Stable Audio Open diffusion model. A small convolutional probe (~125k parameters) is trained to decode frame‑level pitch‑class activations from the model’s latent space using paired audio and MIDI data. During inference, the frozen probe acts as a differentiable loss, guiding generation toward a user‑specified pitch‑class sequence without retraining the base model, and improves melodic coherence by 2.4× over the unguided baseline.

By Yushi Ye, Wilson Zheng, Yongyi Zang
arXiv Machine Learning
1d ago

TrueMuse: A Benchmark for Data Attribution in Text-to-Music Models

TrueMuse is a new benchmark designed to evaluate data attribution in text-to-music models. It consists of a controlled dataset created by fine‑tuning three diffusion‑based models on curated attribution samples, providing known attribution targets. The benchmark covers four settings—melodic structure, timbral characteristics, artist‑level style, and genre‑level patterns—across 133 attributes, 648 models, and 95,456 generated samples, and is used to assess existing black‑box attribution methods along several dimensions.

By Jiawei Yu, Jian Liu
arXiv AI
Sep 15

DuoTok: Source-Aware Dual-Track Music Tokenization for Vocal-Accompaniment Generation

DuoTok is a source‑aware dual‑track music tokenizer designed for vocal‑accompaniment generation. It first learns a semantic audio representation via self‑supervised pretraining, then refines source‑aware structure with feature‑replacement noise and multi‑task supervision (spectral reconstruction, source separation regularization, and an ASR head for lyric alignment). The encoder is frozen and hard‑routed codebooks for vocals and accompaniment are learned, while a diffusion decoder restores fine acoustic detail from the discrete tokens, achieving a favorable predictability‑fidelity trade‑off at ultra‑low bitrate across public benchmarks.

By Rui Lin, Zhiyue Wu, Jiahe Lei, Kangdi Wang, Weixiong Chen, Junyu Dai, Tao Jiang