arXiv Machine Learning

Half-Truth Audio Detection and Localisation: A Lightweight Cross-Attentive Architecture and a Cross-Corpus Diagnostic Study

The paper introduces CAFNet, a lightweight cross‑attentive neural network that fuses MFCC, LFCC, and Chroma‑STFT features to detect and localise partially manipulated (half‑truth) speech. CAFNet achieves high ternary accuracy (97.55%) and low boundary mean absolute error (0.037 s) on the MLADDC benchmark, while demonstrating that cross‑corpus transfer depends on both capability and corpus characteristics. Ablation studies show that cross‑attention fusion is the most critical component, and removing a deeply supervised auxiliary head improves in‑domain performance and reduces variance.

arXiv AI
Aug 17

Teffic-Audio: Tell Fact from Fiction

arXiv:2607. 28351v2 Announce Type: replace-cross Abstract: Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis.

By Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu
arXiv Computation and Language
Sep 24

Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

The paper introduces a trilingual spoken hallucination detection benchmark covering English, Russian, and Kazakh news, with 12,013 samples that include synthetic alterations and severity levels, as well as 290 fact‑checked misinformation items. Detectors are evaluated in a reference‑free setting on text, ASR transcripts, and audio, revealing that most models underperform baseline classifiers, except Gemma‑3n on transcripts. Synthetic‑trained detectors achieve high macro‑F1 scores on real‑world misinformation, but Russian provenance analysis highlights model‑dependent signals that confound synthetic benchmarks.

By Meruyert Aristombayeva, Jason S. Lucas, Chaewan Chun, Dongwon Lee
arXiv AI
Aug 11

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.

By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
arXiv AI
Sep 4

ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection

ToolDF is a tool‑integrated reasoning framework designed for detecting mixed‑authenticity audio deepfakes, where genuine and manipulated audio cues coexist across time or overlapping sources. It uses an audio large language model to orchestrate tasks such as source separation and routing to domain‑specific experts, aggregating their evidence into an interpretable verdict. The authors also introduce a mixed‑authenticity ADD benchmark and report that ToolDF outperforms monolithic baselines, achieving significant macro‑F1 gains while localizing evidence to specific temporal regions and acoustic sources.

By Taewoo Kim, Young Han Lee, Nam In Park, Chanwoo Kim