The paper introduces a domain‑specific parameter‑isolation architecture for domain‑incremental learning (DIL) in audio classification, aiming to preserve knowledge from earlier domains without accessing their data. By employing data‑free generative replay and cross‑domain feature generation, the method constructs new experts conditioned on all previously frozen models, thereby mitigating catastrophic forgetting. Applied to the DCASE 2026 Challenge Task 7, the approach achieves micro and macro accuracies of 78.4 % and 78.9 %, outperforming the baseline by 33 and 25 percentage points, respectively, with ablation studies confirming the contribution of each component.
By Jongyeon Park, Do-Hyeon Lim, Sang-won Park, Hong Kook Kim, Kyungdeuk Ko, Hyeongcheol Geum, Jeong Eun Lim
The paper explores methods to mitigate catastrophic forgetting in incremental learning for sound event classification. It evaluates architectural and regularization strategies on FSD50K and AudioSet, finding that deeper layers, especially the classifier head, are most vulnerable. The most effective approach identified is fully freezing the feature extractor while fine‑tuning a dynamic head, which achieves minimal forgetting, stable training, and a balanced trade‑off between memory stability and learning plasticity.
By Riccardo Casciotti, Annamaria Mesaros
The paper introduces Multiple Embedding Replay Selection (MERS), a graph‑based method that combines supervised and self‑supervised embeddings to improve sample selection for replay buffers in continual learning. MERS replaces traditional buffer selection modules and demonstrates consistent performance gains over state‑of‑the‑art strategies, especially in low‑memory settings. Experiments on CIFAR‑100 and TinyImageNet show that MERS outperforms single‑embedding baselines without adding model parameters or increasing replay volume, making it a practical, drop‑in enhancement for replay‑based continual learning.
By Danit Yanowsky, Daphna Weinshall
arXiv:2607. 03806v1 Announce Type: cross Abstract: Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood.
By H\'ector Martel, Joe Hennessy-Priest, Taemin Cho
arXiv:2608. 19863v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly relied on elaborate pre-training recipes to reach competitive performance.
By Umberto Cappellazzo, Xubo Liu, Stavros Petridis, Maja Pantic
arXiv:2607. 01297v1 Announce Type: cross Abstract: Most existing audio classification methods suppose that each query (testing) sample belongs to a class of support (training) samples, and misrecognize samples of unseen classes as seen classes (cannot reject samples of unseen classes).
By Yanxiong Li, Jiaxin Tan, Qianqian Li, Guoqing Chen, Sen Huang, Tuomas Virtanen
arXiv:2602. 18528v2 Announce Type: replace Abstract: Audio-visual continual test-time adaptation involves continually adapting a source audio-visual model at test-time, to unlabeled non-stationary domains, where either or both modalities can be distributionally shifted, which hampers online cross-modal learning and eventually leads to poor accuracy.
By Sarthak Kumar Maharana, Akshay Mehra, Bhavya Ramakrishna, Yunhui Guo, Guan-Ming Su
Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. While recent CIL advances have spurred significant interest across various modalities, the audio-visual setting remains underexplored.
arXiv:2607. 26607v1 Announce Type: cross Abstract: Few-shot Open-set audio classification requires classifying query samples from known classes with a few labeled support samples while rejecting query samples from unknown classes.
By Tianyan Deng, Yanxiong Li, Rui Gao, Jiahao Du
The paper introduces a generative continual learning framework that uses growing self‑organizing maps (GSOMs) enhanced with learned distributional statistics and encoder‑decoder models for class‑incremental learning. GSOM units store mean, variance, and covariance estimates to synthesize replay samples, enabling exemplar‑free training without raw data or explicit task boundaries. Experiments on multiple benchmarks show that the unsupervised method competes with supervised memory‑based approaches and outperforms memory‑free baselines, especially in single‑class incremental scenarios, and provides baseline results for TinyImageNet and MiniImageNet.
By Pujan Thapa, Alexander Ororbia, Travis Desell
arXiv:2504.10214v2 Announce Type: replace
Abstract: Pretrained model-based incremental object detection (PTMIOD) leverages the rich detection priors of pretrained detectors to learn new categories in...
By Songze Li, Qixing Xu, Tonghua Su, Xu-Yao Zhang, Zhongjie Wang, Yunzhe Li
TEMPO is a unified model that adds temporally‑grounded capabilities to large audio‑language models, enabling timestamping of events, speakers, and sounds in audio, speech, and music. It introduces a supervised fine‑tuning stage featuring atomic timestamp tokens, a time‑aware projector with sinusoidal encodings, and a distance‑aware Gaussian loss, trained via a synthetic‑to‑real curriculum. Additionally, TEMPO employs reinforcement learning (GRPO) as a refinement step, and achieves state‑of‑the‑art performance on a benchmark of 10K samples across five timestamping tasks, surpassing Audio Flamingo Next and Qwen3‑Omni.
By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami, Dinesh Manocha