arXiv:2412. 03771v4 Announce Type: replace-cross Abstract: Zero-shot learning enables models to generalise to unseen classes using semantic information, bridging the gap between training classes and previously unseen test classes.
By Ysobel Sims, Alexandre Mendes, Stephan Chalup
arXiv:2512. 10120v2 Announce Type: replace-cross Abstract: General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting.
By Maris Basha, Anja Zai, Sabine Stoll, Richard Hahnloser
arXiv:2608. 19871v1 Announce Type: new Abstract: Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions by leveraging knowledge of primitive concepts learned from seen compositions.
By Hangyu Tian, Zhenqi He, Yanghao Wang, Long Chen
Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object compositions by leveraging knowledge of primitive concepts learned from seen compositions. Although recent works achieve impressive performance in CZSL by leveraging large vision-language models, they primarily rely on discriminative representations that may not explicitly preserve the structured relationships between primitive concepts and their compositions.
The paper introduces a framework that aligns self‑supervised respiratory encoders with medical terminology in a shared latent space, enabling zero‑shot respiratory sound classification. By using a medical LLM to generate structured reports from metadata, the method creates dense semantic anchors for contrastive learning, combining a sigmoid‑based contrastive loss with the encoder’s native SSL objective and similarity‑aware negative sampling. On nine tasks across six datasets, the approach achieves a 61.3% mean zero‑shot AUC, outperforming CLAP and Qwen2‑Audio, and reaches the highest linear probing AUC with only 43% of the data used by full‑scale baselines.
By Mustafa Talha \.Ilerisoy, Hung Manh Pham, Mathias Funk, Mykola Pechenizkiy, Aaqib Saeed
arXiv:2605. 13672v1 Announce Type: cross Abstract: Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues.
By Giries Abu Ayoub, Morad Tukan, Loay Mualem
SCAPES is a lightweight, resource‑efficient generative model that synthesizes high‑fidelity environmental sounds with high‑level semantic control. It operates on the continuous latent manifold of a neural audio codec, using a segmentation strategy and a Continuous Normalizing Flow to model latent trajectories. A 36‑million‑parameter instance can be trained on limited, uncurated data with a single consumer‑grade GPU, achieving convergence in roughly twice the source audio duration and enabling smooth semantic interpolation.
By Esteban Guti\'errez, Lonce Wyse, Frederic Font, Xavier Serra
arXiv:2607. 15295v1 Announce Type: cross Abstract: We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning.
By Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin
arXiv:2606. 25225v1 Announce Type: cross Abstract: Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning.
By Revant Teotia, Adrien Bardes, Michael Rabbat, Sumit Chopra, Matthew J. Muckley, Nicolas Ballas
arXiv:2608. 04902v1 Announce Type: cross Abstract: Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis.
By Zehua Chen, Junyou Wang, Yuxuan Jiang, Zhenying Fang, Yusheng Dai, Jianfei Chen, Ziwei Liu, Jun Zhu
arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.
By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
PAPT++ is a risk‑aware adversarial generation‑training framework designed to improve single domain generalization. It learns diverse semantic reference images per class and uses them as denoising targets in classifier‑guided diffusion synthesis, thereby generating challenging yet semantically consistent samples. These samples are iteratively combined with source data to update the classifier, progressively exposing it to difficult variations and enhancing generalization performance on standard benchmarks.
By Zhipeng Xu, De Cheng, Xinyang Jiang, Lingfeng He, Huaijie Wang, Dongsheng Li, Nannan Wang, Xinbo Gao