arXiv:2512. 02652v2 Announce Type: replace-cross Abstract: Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance from a symbolic score, rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and language.
By Hong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia, Lan-Zhe Guo, Yu-Feng Li
The paper introduces CPR, a piano rendering framework that combines continuous autoregressive modeling with local flow matching and full‑sequence refinement. It predicts continuous hidden states, generates 24 kHz acoustic latents, and upsamples to 48 kHz, while new techniques BREPA and MT‑RoPE enhance musical semantics and cross‑modal alignment.
By Chong Jing, Junan Zhang, Zhizheng Wu
arXiv:2606. 24307v1 Announce Type: cross Abstract: Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from this domain due to its prohibitive inference latency and offline rendering paradigm.
By Baisen Wang, Chenxi Bao, Qisong Han
SpanSynth-Edit is a flow‑matching model that enables MIDI‑guided synthesis and editing of multi‑instrument audio mixtures using low‑frame‑rate scalar‑quantised latents. It encodes instrument‑labelled note lifecycles as unordered event sets, pools them into a conditioning vector per audio‑latent frame, and uses contextual audio for instrument‑specific timbre guidance. The model supports editing by resynthesising target regions from revised MIDI and demonstrates competitive performance on single‑ and multi‑instrument benchmarks, including within‑frame onset control.
By Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos
arXiv:2607. 27909v1 Announce Type: cross Abstract: Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes.
By Dmitrii Gavrilev, Ilya Borovik, Vladimir Viro
Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performances.