arXiv AI By Marco Pasini, Javier Nistal, Mathias Rose Bjare, Stefan Lattner, George Fazekas

LiveBand: Live Accompaniment Generation in the Audio Domain

Read the original on arXiv AI →

arXiv:2606. 03803v1 Announce Type: cross Abstract: We present LiveBand, a real-time system that generates high-fidelity music accompaniments to live audio input, respecting strict causal constraints.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 8

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

arXiv:2606. 07015v1 Announce Type: cross Abstract: While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the latter overlooks vocal-accompaniment synergy.

By Ziyu Zhang, Chunyu Qiang, Xiaopeng Wang, Yuxin Guo, Kang Yin, Wenjie Tian, Jingbin Hu, Tianlun Zuo, Zhao Guo, Teng Ma, Yuzhe Liang, Chen Zhang, Lei Xie
arXiv AI
Sep 17

CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling

The paper introduces CPR, a piano rendering framework that combines continuous autoregressive modeling with local flow matching and full‑sequence refinement. It predicts continuous hidden states, generates 24 kHz acoustic latents, and upsamples to 48 kHz, while new techniques BREPA and MT‑RoPE enhance musical semantics and cross‑modal alignment.

By Chong Jing, Junan Zhang, Zhizheng Wu
arXiv AI
Sep 15

DuoTok: Source-Aware Dual-Track Music Tokenization for Vocal-Accompaniment Generation

DuoTok is a source‑aware dual‑track music tokenizer designed for vocal‑accompaniment generation. It first learns a semantic audio representation via self‑supervised pretraining, then refines source‑aware structure with feature‑replacement noise and multi‑task supervision (spectral reconstruction, source separation regularization, and an ASR head for lyric alignment). The encoder is frozen and hard‑routed codebooks for vocals and accompaniment are learned, while a diffusion decoder restores fine acoustic detail from the discrete tokens, achieving a favorable predictability‑fidelity trade‑off at ultra‑low bitrate across public benchmarks.

By Rui Lin, Zhiyue Wu, Jiahe Lei, Kangdi Wang, Weixiong Chen, Junyu Dai, Tao Jiang
arXiv Machine Learning
Sep 23

Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning

SpanSynth-Edit is a flow‑matching model that enables MIDI‑guided synthesis and editing of multi‑instrument audio mixtures using low‑frame‑rate scalar‑quantised latents. It encodes instrument‑labelled note lifecycles as unordered event sets, pools them into a conditioning vector per audio‑latent frame, and uses contextual audio for instrument‑specific timbre guidance. The model supports editing by resynthesising target regions from revised MIDI and demonstrates competitive performance on single‑ and multi‑instrument benchmarks, including within‑frame onset control.

By Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos