arXiv AI By Yigitcan \"Ozer, Zhe Zhang, Wanying Ge, Xin Wang, Junichi Yamagishi

A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography

Read the original on arXiv AI →

The paper introduces a training‑free proactive defense for detecting partial deepfake speech by using self‑embedding steganography. It embeds a compressed version of the clean audio within itself, allowing post‑hoc extraction of reference content and enabling detection of spoofed segments via codec‑based restoration. Experiments on a benchmark dataset show that this method complements passive detectors and operates without any training, offering a robust, data‑efficient alternative for partial deepfake detection.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 11

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.

By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
arXiv Machine Learning
Sep 4

CRAW: Codec Robust Audio Watermarking

The paper introduces CRAW, a codec‑robust audio watermarking framework designed to embed imperceptible signals into synthetic speech. CRAW enhances robustness against neural re‑synthesis, codecs, denoisers, and vocoders while preserving high perceptual quality through distortion‑aware training, attention‑based pooling, perceptual masking, and error‑correcting codes. Experiments show CRAW outperforms existing post‑hoc watermarking methods in robustness without compromising audio quality.

By David Chernin, Ethan Fetaya
Hugging Face Trending Papers
Sep 2

CRAW: Codec Robust Audio Watermarking

CRAW is a codec‑robust audio watermarking framework designed to embed imperceptible signals into synthetic speech, enabling provenance verification. It improves robustness against neural re‑synthesis, codecs, denoisers, and vocoders while preserving perceptual quality through distortion‑aware training, attention‑based pooling, perceptual masking, and error‑correcting codes. Experiments show CRAW outperforms existing post‑hoc watermarking methods in robustness without compromising audio quality.