arXiv:2606. 10912v1 Announce Type: cross Abstract: Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence lies, or what cues drive the decision.
By Vojt\v{e}ch Stan\v{e}k, Veronika Jirmusov\'a, Anton Firc, Kamil Malinka, Jakub Re\v{s}, Martin Pere\v{s}\'ini
ToolDF is a tool‑integrated reasoning framework designed for detecting mixed‑authenticity audio deepfakes, where genuine and manipulated audio cues coexist across time or overlapping sources. It uses an audio large language model to orchestrate tasks such as source separation and routing to domain‑specific experts, aggregating their evidence into an interpretable verdict. The authors also introduce a mixed‑authenticity ADD benchmark and report that ToolDF outperforms monolithic baselines, achieving significant macro‑F1 gains while localizing evidence to specific temporal regions and acoustic sources.
By Taewoo Kim, Young Han Lee, Nam In Park, Chanwoo Kim
arXiv:2606. 16532v1 Announce Type: cross Abstract: Audio deepfake detectors often fail to generalize across speakers, as they learn speaker-identity features rather than synthesis artifacts, known as implicit identity leakage.
By Zhuodong Liu, Hugen Lv, Xiangyu Li, Chunhong Yuan
arXiv:2608.23363v1 Announce Type: cross
Abstract: Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deep...
By Vlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu, A. Sophia Koepke, Radu Tudor Ionescu
arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.
By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
arXiv:2603. 23667v2 Announce Type: replace-cross Abstract: We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions.
By Octavian Pascu, Dan Oneata, Horia Cucu, Nicolas M. Muller
arXiv:2607. 04848v1 Announce Type: cross Abstract: While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects.
By Linxi Li, Yuncong Yu, Qianwei Guo, Liwei Jin, Yechen Wang, Carsten Maple
arXiv:2510.12851v2 Announce Type: replace-cross
Abstract: Large Audio-Language Models (LALMs) excel in Audio QA but often suffer from hallucinations ungrounded in the audio. To our knowledge, we are...
By Tsung-En Lin, Kuan-Yi Lee, Hung-Yi Lee
arXiv:2609.37586v1 Announce Type: cross
Abstract: Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech....
By Yuankun Xie, Xiaoxuan Guo, Xiaopeng Wang, Siqing Qin, Shaole Li, Kong Aik Lee
arXiv:2607. 28351v2 Announce Type: replace-cross Abstract: Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis.
By Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu
arXiv:2603. 01006v3 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth.
By Pengfei Zhang, Tianxin Xie, Minghao Yang, Li Liu
arXiv:2606. 05101v1 Announce Type: cross Abstract: Audio deepfake detection (ADD) models are critical for countering the malicious use of text-to-speech (TTS) models.
By Sepehr Dehdashtian, Jacob H Seidman, Vishnu N Boddeti, Gaurav Bharaj