arXiv:2606. 09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling.
By Wei Fan, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Kejiang Chen, Weiming Zhang, Nenghai Yu
arXiv:2606. 30700v1 Announce Type: cross Abstract: Self-supervised learning enables audio representations that transfer across domains and tasks.
By Ludovic K. Tuncay (IRIT-SAMoVA), Etienne Labb\'e (IRIT-SAMoVA), Thomas Pellegrini (IRIT-SAMoVA)
The paper investigates how procedural audio should be scaled for effective pre‑training and whether training strategies from natural audio transfer to procedural data. By separating scale into formula‑class coverage and within‑class rendering diversity, the authors show that each type of scale benefits different learning formulations and downstream tasks. Their experiments reveal that procedural audio prefers lower mask ratios, and that it exhibits lower patch diversity and stronger temporal predictability compared to natural audio, leading to a proposal for source‑aware procedural pre‑training.
By Jiajun Peng, Fengrui Liu, Xinyu Liu, Feng Liu
arXiv:2606. 06357v1 Announce Type: cross Abstract: Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but are not directly decodable.
By Dinghao Zhou, Xingchen Song, Di Wu, Pengyu Cheng, Shengfan Shen, Sixiang Lv
arXiv:2606. 30356v1 Announce Type: cross Abstract: We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives.
By Karl El Hajal, Mathew Magimai. -Doss
SCAPES is a lightweight, resource‑efficient generative model that synthesizes high‑fidelity environmental sounds with high‑level semantic control. It operates on the continuous latent manifold of a neural audio codec, using a segmentation strategy and a Continuous Normalizing Flow to model latent trajectories. A 36‑million‑parameter instance can be trained on limited, uncurated data with a single consumer‑grade GPU, achieving convergence in roughly twice the source audio duration and enabling smooth semantic interpolation.
By Esteban Guti\'errez, Lonce Wyse, Frederic Font, Xavier Serra