arXiv Machine Learning

Architecture and Affordances of PLAUD: Performative Latents and Unsupervised DDSP

arXiv:2608. 13724v1 Announce Type: cross Abstract: PLAUD (Performative Latents and Unsupervised DDSP) is a neural synthesizer and Max for Live instrument for live electronic music, built on NoiseBandNet and trained on small personal sound corpora.

arXiv AI
Jun 26

Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training

arXiv:2512. 02652v2 Announce Type: replace-cross Abstract: Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance from a symbolic score, rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and language.

By Hong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia, Lan-Zhe Guo, Yu-Feng Li
arXiv AI
Sep 7

Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes

The paper introduces a lightweight technique to steer the pitch content of audio generated by the Stable Audio Open diffusion model. A small convolutional probe (~125k parameters) is trained to decode frame‑level pitch‑class activations from the model’s latent space using paired audio and MIDI data. During inference, the frozen probe acts as a differentiable loss, guiding generation toward a user‑specified pitch‑class sequence without retraining the base model, and improves melodic coherence by 2.4× over the unguided baseline.

By Yushi Ye, Wilson Zheng, Yongyi Zang
arXiv Machine Learning
Sep 23

Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning

SpanSynth-Edit is a flow‑matching model that enables MIDI‑guided synthesis and editing of multi‑instrument audio mixtures using low‑frame‑rate scalar‑quantised latents. It encodes instrument‑labelled note lifecycles as unordered event sets, pools them into a conditioning vector per audio‑latent frame, and uses contextual audio for instrument‑specific timbre guidance. The model supports editing by resynthesising target regions from revised MIDI and demonstrates competitive performance on single‑ and multi‑instrument benchmarks, including within‑frame onset control.

By Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos
arXiv Machine Learning
5d ago

Synth-JEPA: Joint Embedding Prediction for Renderer-Free Synthesizer Parameter Search

Synth-JEPA introduces a renderer‑free approach to synthesizer parameter search by learning mutually predictive audio and parameter representations from paired synthesizer data. During inference, candidate parameters are scored directly in this learned space, avoiding the need to render each candidate and shaping audio geometry through parameter correspondences. Evaluations on Surge XT and out‑of‑domain datasets show Synth‑JEPA outperforms inverse models, direct search, and learned proxy objectives, with listeners preferring its matches in 85% of pairwise tests.

By Ben Hayes, Haokun Tian, Stefan Lattner