arXiv AI

RAT: Reference-Augmented Training for ASV Anti-Spoofing

arXiv:2606. 10908v1 Announce Type: cross Abstract: We introduce a spoofing countermeasure architecture conditioned on speaker-reference recordings, but observe that it converges to a solution that effectively ignores the reference during inference.

arXiv AI
Aug 17

Teffic-Audio: Tell Fact from Fiction

arXiv:2607. 28351v2 Announce Type: replace-cross Abstract: Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis.

By Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu
arXiv AI
Aug 18

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

arXiv:2608. 15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output.

By Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov
arXiv AI
Sep 21

CoReLoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection

CoReLoop introduces a parameter‑efficient refinement strategy for audio deepfake detection that reuses a frozen SSL‑based detector’s encoder outputs without altering its original parameters. By adapting recurrent inputs, controlling state updates, and aligning refined outputs with the frozen classifier, the method adds lightweight refinement modules and low‑rank adapters trained on the original data. On 14 cross‑domain test sets, the 24‑layer model reduces pooled equal error rate from 4.85% to 3.74% with two passes, and an optional halting head further improves performance to 3.73% with an average of 1.18 passes.

By Kunyu Feng, Yuxiang Wang, Li Wang, Wan Lin, Zhizheng Wu