Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem, continual learning techniques are nowadays widely considered feasible strategies for updating models to new datasets, but they also lead to decreased performance on previously seen samples (catastrophic forgetting).
arXiv:2606. 14459v1 Announce Type: cross Abstract: Modern Automatic Speech Recognition (ASR) systems have made remarkable progress on standard benchmarks, yet performance gaps have emerged under real-world distribution shifts, caused by recording conditions, accents, speech impairments, and noise.
By Theresa Pekarek Rosin, Matthias Kerzel, Stefan Wermter
The paper introduces a domain‑specific parameter‑isolation architecture for domain‑incremental learning (DIL) in audio classification, aiming to preserve knowledge from earlier domains without accessing their data. By employing data‑free generative replay and cross‑domain feature generation, the method constructs new experts conditioned on all previously frozen models, thereby mitigating catastrophic forgetting. Applied to the DCASE 2026 Challenge Task 7, the approach achieves micro and macro accuracies of 78.4 % and 78.9 %, outperforming the baseline by 33 and 25 percentage points, respectively, with ablation studies confirming the contribution of each component.
By Jongyeon Park, Do-Hyeon Lim, Sang-won Park, Hong Kook Kim, Kyungdeuk Ko, Hyeongcheol Geum, Jeong Eun Lim
arXiv:2609.37586v1 Announce Type: cross
Abstract: Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech....
By Yuankun Xie, Xiaoxuan Guo, Xiaopeng Wang, Siqing Qin, Shaole Li, Kong Aik Lee
The paper explores methods to mitigate catastrophic forgetting in incremental learning for sound event classification. It evaluates architectural and regularization strategies on FSD50K and AudioSet, finding that deeper layers, especially the classifier head, are most vulnerable. The most effective approach identified is fully freezing the feature extractor while fine‑tuning a dynamic head, which achieves minimal forgetting, stable training, and a balanced trade‑off between memory stability and learning plasticity.
By Riccardo Casciotti, Annamaria Mesaros
arXiv:2606. 14391v1 Announce Type: cross Abstract: Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are often optimized to omit disfluencies, leading to information loss and hallucinations.
By Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee, Stefan Wermter