arXiv Machine Learning By Ysobel Sims, Alexandre Mendes, Stephan Chalup

A Benchmark of Generative Methods for Zero-Shot Environmental Sound Classification

Read the original on arXiv Machine Learning →

arXiv:2412. 03771v4 Announce Type: replace-cross Abstract: Zero-shot learning enables models to generalise to unseen classes using semantic information, bridging the gap between training classes and previously unseen test classes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 7

SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds

SCAPES is a lightweight, resource‑efficient generative model that synthesizes high‑fidelity environmental sounds with high‑level semantic control. It operates on the continuous latent manifold of a neural audio codec, using a segmentation strategy and a Continuous Normalizing Flow to model latent trajectories. A 36‑million‑parameter instance can be trained on limited, uncurated data with a single consumer‑grade GPU, achieving convergence in roughly twice the source audio duration and enabling smooth semantic interpolation.

By Esteban Guti\'errez, Lonce Wyse, Frederic Font, Xavier Serra