arXiv:2606. 31587v1 Announce Type: cross Abstract: Audio-Language Models (ALMs) achieve strong zero-shot performance by aligning audio with textual class descriptions.
By Asif Hanif, Mohammad Yaqub
arXiv:2606. 08898v1 Announce Type: cross Abstract: In the task of few-shot class-incremental audio classification, the number of classes is assumed to always increase without considering the possibility of decrease.
By Yanxiong Li, Guoqing Chen, Qianqian Li, Sen Huang
arXiv:2606. 03843v1 Announce Type: cross Abstract: Continual learning methods aim to maximize the stability and plasticity of machine learning models that are trained on a sequence of tasks.
By Amogh Inamdar, Matthew So, Vici Milenia, Richard Zemel
arXiv:2607. 01297v1 Announce Type: cross Abstract: Most existing audio classification methods suppose that each query (testing) sample belongs to a class of support (training) samples, and misrecognize samples of unseen classes as seen classes (cannot reject samples of unseen classes).
By Yanxiong Li, Jiaxin Tan, Qianqian Li, Guoqing Chen, Sen Huang, Tuomas Virtanen
arXiv:2606. 14459v1 Announce Type: cross Abstract: Modern Automatic Speech Recognition (ASR) systems have made remarkable progress on standard benchmarks, yet performance gaps have emerged under real-world distribution shifts, caused by recording conditions, accents, speech impairments, and noise.
By Theresa Pekarek Rosin, Matthias Kerzel, Stefan Wermter
arXiv:2502. 16584v2 Announce Type: replace-cross Abstract: Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs).
By Liumeng Xue, Ziya Zhou, Jiahao Pan, Zixuan Li, Shuai Fan, Yinghao Ma, Sitong Cheng, Dongchao Yang, Haohan Guo, Yujia Xiao, Xinsheng Wang, Zixuan Shen, Chuanbo Zhu, Xinshen Zhang, Tianchi Liu, Ruibin Yuan, Zeyue Tian, Haohe Liu, Xingjian Du, Emmanouil Benetos, Ge Zhang, Yike Guo, Wei Xue
arXiv:2606. 20518v1 Announce Type: new Abstract: Flow-matching text-to-speech systems achieve remarkable zero-shot quality but remain static after deployment: pronunciation errors on out-of-vocabulary proper nouns persist unless the model is retrained.
By Harshit Singh, Ayush Pratap Singh, Nityanand Mathur
arXiv:2606. 02615v1 Announce Type: cross Abstract: Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition.
By Haolong Zheng, Siyin Wang, Xulin Fan, Zengrui Jin, Mark Hasegawa-Johnson
arXiv:2606.02615v2 Announce Type: replace-cross
Abstract: Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognit...
By Haolong Zheng, Siyin Wang, Xulin Fan, Zengrui Jin, Mark Hasegawa-Johnson
The paper explores methods to mitigate catastrophic forgetting in incremental learning for sound event classification. It evaluates architectural and regularization strategies on FSD50K and AudioSet, finding that deeper layers, especially the classifier head, are most vulnerable. The most effective approach identified is fully freezing the feature extractor while fine‑tuning a dynamic head, which achieves minimal forgetting, stable training, and a balanced trade‑off between memory stability and learning plasticity.
By Riccardo Casciotti, Annamaria Mesaros
The paper introduces PromptCCZSL, a framework that enables vision‑language models to continually learn new attributes, objects, and their unique compositions while avoiding forgetting. It uses a frozen VLM backbone with prompt‑based techniques, recency‑weighted multi‑teacher distillation, and several loss functions (CAL, OPL, IDL) to maintain prior knowledge and promote diverse, distinct embeddings. Experiments on UT‑Zappos and C‑GQA show significant performance gains over existing VLM‑based and non‑VLM baselines, establishing a new benchmark for continual compositional zero‑shot learning.
By Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.