arXiv AI By Yushi Ye, Wilson Zheng, Yongyi Zang

Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes

Read the original on arXiv AI →

The paper introduces a lightweight technique to steer the pitch content of audio generated by the Stable Audio Open diffusion model. A small convolutional probe (~125k parameters) is trained to decode frame‑level pitch‑class activations from the model’s latent space using paired audio and MIDI data. During inference, the frozen probe acts as a differentiable loss, guiding generation toward a user‑specified pitch‑class sequence without retraining the base model, and improves melodic coherence by 2.4× over the unguided baseline.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 8

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

arXiv:2606. 07015v1 Announce Type: cross Abstract: While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the latter overlooks vocal-accompaniment synergy.

By Ziyu Zhang, Chunyu Qiang, Xiaopeng Wang, Yuxin Guo, Kang Yin, Wenjie Tian, Jingbin Hu, Tianlun Zuo, Zhao Guo, Teng Ma, Yuzhe Liang, Chen Zhang, Lei Xie
arXiv AI
Sep 1

PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation

PhysWave is a physics-guided latent diffusion model designed for controllable text-to-First-Order Ambisonics (FOA) audio generation. It integrates natural-language and trajectory control via a shared waypoint-caption representation and incorporates two differentiable acoustic priors—spherical-harmonic direction consistency and inverse-square distance consistency—into diffusion training. The authors also introduce a 300K-clip FOA dataset and demonstrate that these priors improve spatial consistency while preserving audio quality, with potential use as inference-time guidance for training-free refinement.

By Lingfeng Yao, Chenpei Huang, Xingke Yang, Ziye Geng, Changqing Luo, Hao Wang, Jiang Liu, Miao Pan