From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation
arXiv:2606. 14791v1 Announce Type: cross Abstract: Self-supervised learning advances audio representation for multimedia analysis.
The paper investigates how procedural audio should be scaled for effective pre‑training and whether training strategies from natural audio transfer to procedural data. By separating scale into formula‑class coverage and within‑class rendering diversity, the authors show that each type of scale benefits different learning formulations and downstream tasks. Their experiments reveal that procedural audio prefers lower mask ratios, and that it exhibits lower patch diversity and stronger temporal predictability compared to natural audio, leading to a proposal for source‑aware procedural pre‑training.
arXiv:2606. 14791v1 Announce Type: cross Abstract: Self-supervised learning advances audio representation for multimedia analysis.
SsgCaps is a publicly available dataset of human-engineered sound scenes, each paired with a precisely structured prompt that guides the sampling process. The prompts are drawn from a predefined action-based typology, enabling extensive yet plausible sampling. A comparative quantitative analysis shows only small differences between the open and private versions, supporting the recommendation of the open version for benchmarking sound scene generation algorithms.
arXiv:2607. 09973v1 Announce Type: cross Abstract: Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows.
arXiv:2606. 30700v1 Announce Type: cross Abstract: Self-supervised learning enables audio representations that transfer across domains and tasks.
arXiv:2609.17076v1 Announce Type: new Abstract: Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting represe...
arXiv:2511. 16757v2 Announce Type: replace-cross Abstract: Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored.
arXiv:2506. 20995v4 Announce Type: replace-cross Abstract: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis.
arXiv:2606. 00629v1 Announce Type: cross Abstract: Sound design workflows frequently oscillate between time-consuming library searches and the complexity of procedural synthesis, with practitioners typically relying on disconnected tools to address each challenge separately.
arXiv:2601. 18904v3 Announce Type: replace-cross Abstract: Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data.
arXiv:2603. 01006v3 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth.
arXiv:2606. 10147v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to shape an answer?
arXiv:2608. 19863v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly relied on elaborate pre-training recipes to reach competitive performance.