Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. We use a Diffusion Transformer (DiT), trained only on bona fide speech, as a frozen reconstruction probe.
arXiv:2609.08899v2 Announce Type: replace-cross
Abstract: Speech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. This has driven strong progress...
By Mengzhe Geng, Yujia Lu, Patrick Littell, Manuela Kunz, Xie Chen
arXiv:2511. 21325v2 Announce Type: replace-cross Abstract: Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs.
By Ido Nitzan Hidekel, Gal lifshitz, Khen Cohen, Dan Raviv
The paper introduces NEUROTOKEN, a unified neural network for auditory attention decoding (AAD) that jointly predicts the attended speaker’s direction and source by modeling the conditional likelihood of the attended envelope given EEG. It employs a conditional flow‑matching head (ATTUNEFLOW) and two inference‑time ensembles (QUADTRACK and ENV‑FLOW) to improve source‑AAD accuracy and reduce variance across subjects. Experiments on KU Leuven, DTU, and NJU datasets show significant gains over existing baselines and reveal that prior direction‑AAD results overestimate performance under stricter protocols.
By Ali Alavi, Donald S. Williamson
The paper introduces CAFNet, a lightweight cross‑attentive neural network that fuses MFCC, LFCC, and Chroma‑STFT features to detect and localise partially manipulated (half‑truth) speech. CAFNet achieves high ternary accuracy (97.55%) and low boundary mean absolute error (0.037 s) on the MLADDC benchmark, while demonstrating that cross‑corpus transfer depends on both capability and corpus characteristics. Ablation studies show that cross‑attention fusion is the most critical component, and removing a deeply supervised auxiliary head improves in‑domain performance and reduces variance.
By S. Sutharya, Remya K. Sasi
arXiv:2609.13842v1 Announce Type: cross
Abstract: Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing a...
By Minh-Xuan Phan, Khalid Zaman, Candy Olivia Mawalim, Masashi Unoki