The paper explores methods to mitigate catastrophic forgetting in incremental learning for sound event classification. It evaluates architectural and regularization strategies on FSD50K and AudioSet, finding that deeper layers, especially the classifier head, are most vulnerable. The most effective approach identified is fully freezing the feature extractor while fine‑tuning a dynamic head, which achieves minimal forgetting, stable training, and a balanced trade‑off between memory stability and learning plasticity.
By Riccardo Casciotti, Annamaria Mesaros
Fully few-shot class-incremental audio classification (FFCAC) requires recognizing new sound classes from only a handful of labeled examples per session, without forgetting previously learned classes...
arXiv:2602. 18528v2 Announce Type: replace Abstract: Audio-visual continual test-time adaptation involves continually adapting a source audio-visual model at test-time, to unlabeled non-stationary domains, where either or both modalities can be distributionally shifted, which hampers online cross-modal learning and eventually leads to poor accuracy.
By Sarthak Kumar Maharana, Akshay Mehra, Bhavya Ramakrishna, Yunhui Guo, Guan-Ming Su
arXiv:2609.37586v1 Announce Type: cross
Abstract: Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech....
By Yuankun Xie, Xiaoxuan Guo, Xiaopeng Wang, Siqing Qin, Shaole Li, Kong Aik Lee
The paper introduces BTS-CAFE, a federated domain generalization framework for respiratory sound classification that addresses stethoscope-induced shortcuts. It combines causality-inspired device-style interventions, counterfactual metadata augmentation, and gradient alignment to reduce style–content entanglement and promote device-invariant decision boundaries. Experiments on ICBHI and SPRSound datasets show a 3.69‑point improvement in out-of-distribution performance over the baseline and outperform conventional data augmentation and federated learning methods.
By Heejoon Koo, Yoon Tae Kim, Miika Toikkanen, June-Woo Kim
arXiv:2607. 12569v1 Announce Type: cross Abstract: Fake speech detectors are increasingly challenged by the development of new and more accurate generative models.
By Enrico Gottardis, Mattia Tamiazzo, Simone Milani
The paper introduces HiDeR, a High Information Density Replay framework for Lifelong Person Re-Identification that replaces discrete sample selection with information compression. It uses a complexity‑aware memory allocation based on intra‑class variance and a metric‑guided condensation objective to preserve essential identity topologies, while a cross‑modality adaptation strategy bridges synthetic and real styles to improve training. Experiments show HiDeR outperforms state‑of‑the‑art methods in knowledge retention and generalization, and reduces cumulative replay cost.
By Mingyu Wang, Wei Jiang, Haojie Liu, Zhiyong Li, Weijie Mao
arXiv:2607. 20493v1 Announce Type: new Abstract: Deep learning has led to remarkable progress in artificial intelligence, particularly in robotics, imaging and sound processing.
By Quentin Besnard (RFAI), Nicolas Ragot (RFAI)
The paper investigates how streaming emotion recognition models can be misled by their own prior predictions, a problem termed previous-belief contamination (PBC). Using a counterfactual diagnostic on CREMA-D-Stream, the authors show that feeding a model’s previous emotion label into its current prediction can drastically lower accuracy and flip many predictions, with the effect varying by label. To mitigate PBC, they propose EmoUpdate, a training‑free framework that isolates current audio perception from historical context through a prior‑blind firewall, a causal belief filter, and a decontamination operator, achieving significant gains across multiple SpeechLMs and benchmarks.
By Haoyue Liu, Zhichao Wang, Ye Chen, Haonan Deng, Xiaoying Tang
Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. While recent CIL advances have spurred significant interest across various modalities, the audio-visual setting remains underexplored.
arXiv:2608. 19863v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly relied on elaborate pre-training recipes to reach competitive performance.
By Umberto Cappellazzo, Xubo Liu, Stavros Petridis, Maja Pantic
arXiv:2609.17981v1 Announce Type: cross
Abstract: Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-R...
By Mohan Shi, Zilai Wang, Natarajan Balaji Shankar, Kaiyuan Zhang, Eray Eren, Abeer Alwan