Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning
Read the original on arXiv Machine Learning →SpanSynth-Edit is a flow‑matching model that enables MIDI‑guided synthesis and editing of multi‑instrument audio mixtures using low‑frame‑rate scalar‑quantised latents. It encodes instrument‑labelled note lifecycles as unordered event sets, pools them into a conditioning vector per audio‑latent frame, and uses contextual audio for instrument‑specific timbre guidance. The model supports editing by resynthesising target regions from revised MIDI and demonstrates competitive performance on single‑ and multi‑instrument benchmarks, including within‑frame onset control.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.