The paper introduces a domain‑specific parameter‑isolation architecture for domain‑incremental learning (DIL) in audio classification, aiming to preserve knowledge from earlier domains without accessing their data. By employing data‑free generative replay and cross‑domain feature generation, the method constructs new experts conditioned on all previously frozen models, thereby mitigating catastrophic forgetting. Applied to the DCASE 2026 Challenge Task 7, the approach achieves micro and macro accuracies of 78.4 % and 78.9 %, outperforming the baseline by 33 and 25 percentage points, respectively, with ablation studies confirming the contribution of each component.
By Jongyeon Park, Do-Hyeon Lim, Sang-won Park, Hong Kook Kim, Kyungdeuk Ko, Hyeongcheol Geum, Jeong Eun Lim
CoReLoop introduces a parameter‑efficient refinement strategy for audio deepfake detection that reuses a frozen SSL‑based detector’s encoder outputs without altering its original parameters. By adapting recurrent inputs, controlling state updates, and aligning refined outputs with the frozen classifier, the method adds lightweight refinement modules and low‑rank adapters trained on the original data. On 14 cross‑domain test sets, the 24‑layer model reduces pooled equal error rate from 4.85% to 3.74% with two passes, and an optional halting head further improves performance to 3.73% with an average of 1.18 passes.
By Kunyu Feng, Yuxiang Wang, Li Wang, Wan Lin, Zhizheng Wu
CoRELoop introduces a parameter‑efficient refinement framework for audio deepfake detection that operates on a pre‑trained SSL‑based detector without altering its original parameters. By adapting recurrent inputs, controlling state updates, and aligning refined outputs with the frozen classifier, CoRELoop adds lightweight refinement modules and low‑rank adapters, achieving a pooled equal error rate reduction from 4.85% to 3.74% on 14 cross‑domain test sets with only about 10 M trainable parameters. An optional halting head further optimizes performance, reaching 3.73% pooled EER with an average of 1.18 passes.
By Kunyu Feng, Yuxiang Wang, Li Wang, Wan Lin, Zhizheng Wu
arXiv:2607. 17761v1 Announce Type: cross Abstract: Recently, speech deepfake detection (SDD) has achieved significant progress.
By Jun Xue, Zhuolin Yi, Yanzhen Ren, Yihuan Huang, Jiayu Xiong, Yi Chai, Guanxiang Feng, Jiajun Liu, Tong Zhang
arXiv:2608. 13817v1 Announce Type: cross Abstract: Human speech production is constrained by physiology, giving rise to characteristic temporal structure on acoustic signals.
By Tom\'as Andrade Weber
arXiv:2606. 14459v1 Announce Type: cross Abstract: Modern Automatic Speech Recognition (ASR) systems have made remarkable progress on standard benchmarks, yet performance gaps have emerged under real-world distribution shifts, caused by recording conditions, accents, speech impairments, and noise.
By Theresa Pekarek Rosin, Matthias Kerzel, Stefan Wermter
arXiv:2606. 14391v1 Announce Type: cross Abstract: Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are often optimized to omit disfluencies, leading to information loss and hallucinations.
By Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee, Stefan Wermter
arXiv:2606. 05121v1 Announce Type: cross Abstract: Audio is an inherently interactive modality, yet today's Large Audio Language Models (LALMs) are offline, and streaming audio models each handle only a single task such as streaming ASR or voice chatting.
By Zhifei Xie, Zihang Liu, Ze An, Xiaobin Hu, Yue Liao, Ziyang Ma, Dongchao Yang, Mingbao Lin, Deheng Ye, Shuicheng Yan, Chunyan Miao
arXiv:2606. 10912v1 Announce Type: cross Abstract: Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence lies, or what cues drive the decision.
By Vojt\v{e}ch Stan\v{e}k, Veronika Jirmusov\'a, Anton Firc, Kamil Malinka, Jakub Re\v{s}, Martin Pere\v{s}\'ini
arXiv:2607. 12569v1 Announce Type: cross Abstract: Fake speech detectors are increasingly challenged by the development of new and more accurate generative models.
By Enrico Gottardis, Mattia Tamiazzo, Simone Milani
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
By Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin
arXiv:2606. 19579v1 Announce Type: cross Abstract: Audio deepfakes generated by neural text-to-speech and voice-cloning systems threaten speaker verification and public discourse at scale.
By Shivaay Dhondiyal, Divyansh Sharma, Dinesh Kumar Vishwakarma