Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces a domain‑specific parameter‑isolation architecture for domain‑incremental learning (DIL) in audio classification, aiming to preserve knowledge from earlier domains without accessing their data. By employing data‑free generative replay and cross‑domain feature generation, the method constructs new experts conditioned on all previously frozen models, thereby mitigating catastrophic forgetting. Applied to the DCASE 2026 Challenge Task 7, the approach achieves micro and macro accuracies of 78.4 % and 78.9 %, outperforming the baseline by 33 and 25 percentage points, respectively, with ablation studies confirming the contribution of each component.
CoReLoop introduces a parameter‑efficient refinement strategy for audio deepfake detection that reuses a frozen SSL‑based detector’s encoder outputs without altering its original parameters. By adapting recurrent inputs, controlling state updates, and aligning refined outputs with the frozen classifier, the method adds lightweight refinement modules and low‑rank adapters trained on the original data. On 14 cross‑domain test sets, the 24‑layer model reduces pooled equal error rate from 4.85% to 3.74% with two passes, and an optional halting head further improves performance to 3.73% with an average of 1.18 passes.
CoRELoop introduces a parameter‑efficient refinement framework for audio deepfake detection that operates on a pre‑trained SSL‑based detector without altering its original parameters. By adapting recurrent inputs, controlling state updates, and aligning refined outputs with the frozen classifier, CoRELoop adds lightweight refinement modules and low‑rank adapters, achieving a pooled equal error rate reduction from 4.85% to 3.74% on 14 cross‑domain test sets with only about 10 M trainable parameters. An optional halting head further optimizes performance, reaching 3.73% pooled EER with an average of 1.18 passes.
arXiv:2607. 17761v1 Announce Type: cross Abstract: Recently, speech deepfake detection (SDD) has achieved significant progress.
arXiv:2608. 13817v1 Announce Type: cross Abstract: Human speech production is constrained by physiology, giving rise to characteristic temporal structure on acoustic signals.
arXiv:2606. 14459v1 Announce Type: cross Abstract: Modern Automatic Speech Recognition (ASR) systems have made remarkable progress on standard benchmarks, yet performance gaps have emerged under real-world distribution shifts, caused by recording conditions, accents, speech impairments, and noise.