The paper introduces a lightweight technique to steer the pitch content of audio generated by the Stable Audio Open diffusion model. A small convolutional probe (~125k parameters) is trained to decode frame‑level pitch‑class activations from the model’s latent space using paired audio and MIDI data. During inference, the frozen probe acts as a differentiable loss, guiding generation toward a user‑specified pitch‑class sequence without retraining the base model, and improves melodic coherence by 2.4× over the unguided baseline.
By Yushi Ye, Wilson Zheng, Yongyi Zang
arXiv:2609.39552v1 Announce Type: cross
Abstract: Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these pheno...
By Arhan Vohra, Choenden Kyirong, Laura Ib\'a\~nez-Mart\'inez, Mart\'in Rocamora
arXiv:2608. 03920v1 Announce Type: cross Abstract: Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space.
By Zixun Guo, Simon Dixon
arXiv:2606. 04040v1 Announce Type: cross Abstract: Brain-computer interfaces aim to decode naturalistic stimuli from neural signals, yet most progress to date has focused on vision and language.
By Jiaxin Qing, Junwei Lu, Lexin Li
arXiv:2607. 03806v1 Announce Type: cross Abstract: Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood.
By H\'ector Martel, Joe Hennessy-Priest, Taemin Cho
arXiv:2606. 06357v1 Announce Type: cross Abstract: Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but are not directly decodable.
By Dinghao Zhou, Xingchen Song, Di Wu, Pengyu Cheng, Shengfan Shen, Sixiang Lv