arXiv Machine Learning By Fengrui Liu, Ruiyang Huang, Qijian Zheng, Yuanfang Wang, Feng Liu

From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation

Read the original on arXiv Machine Learning →

arXiv:2606. 14791v1 Announce Type: cross Abstract: Self-supervised learning advances audio representation for multimedia analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 15

Rethinking Procedural Audio Pre-training: Source Scaling and Objective Adaptation

The paper investigates how procedural audio should be scaled for effective pre‑training and whether training strategies from natural audio transfer to procedural data. By separating scale into formula‑class coverage and within‑class rendering diversity, the authors show that each type of scale benefits different learning formulations and downstream tasks. Their experiments reveal that procedural audio prefers lower mask ratios, and that it exhibits lower patch diversity and stronger temporal predictability compared to natural audio, leading to a proposal for source‑aware procedural pre‑training.

By Jiajun Peng, Fengrui Liu, Xinyu Liu, Feng Liu
arXiv AI
Sep 7

SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds

SCAPES is a lightweight, resource‑efficient generative model that synthesizes high‑fidelity environmental sounds with high‑level semantic control. It operates on the continuous latent manifold of a neural audio codec, using a segmentation strategy and a Continuous Normalizing Flow to model latent trajectories. A 36‑million‑parameter instance can be trained on limited, uncurated data with a single consumer‑grade GPU, achieving convergence in roughly twice the source audio duration and enabling smooth semantic interpolation.

By Esteban Guti\'errez, Lonce Wyse, Frederic Font, Xavier Serra