arXiv AI By Kyudan Jung, Jihwan Kim, Minwoo Lee, Soyoon Kim, Jeonghoon Kim, Jaegul Choo, Cheonbok Park

SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection

Read the original on arXiv AI →

The paper introduces SNAP, a speaker‑nulling framework designed to improve deepfake speech detection. By estimating a speaker subspace and orthogonally projecting out speaker‑dependent components, SNAP isolates synthesis artifacts in the residual features. This reduction of speaker entanglement enables detectors to focus on artifact‑related cues, achieving state‑of‑the‑art performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 16

Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

The paper proposes a method called language orthogonalization to improve zero‑shot cross‑lingual audio deepfake detection. By removing language‑dependent variation from self‑supervised speech models using a target‑free ridge map on language‑identification embeddings, the approach consistently lowers equal error rates across six languages and six model backbones. The gains are larger when the target language is more distant in the language‑identification space.

By Minu Kim, Ji Sub Um, Hoirin Kim
arXiv AI
Aug 11

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.

By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang