Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models such as Chatterbox, provided through dedicated techniques such as AudioSeal, or deployed by commercial platforms such as ElevenLabs. We identify a previously uncharacterized liability: when synthetic speech is watermarked and human speech is not, detectors trained alongside latch onto the watermark as a spurious "watermark => fake" shortcut.
The paper introduces CRAW, a codec‑robust audio watermarking framework designed to embed imperceptible signals into synthetic speech. CRAW enhances robustness against neural re‑synthesis, codecs, denoisers, and vocoders while preserving high perceptual quality through distortion‑aware training, attention‑based pooling, perceptual masking, and error‑correcting codes. Experiments show CRAW outperforms existing post‑hoc watermarking methods in robustness without compromising audio quality.
By David Chernin, Ethan Fetaya
The paper introduces Traceable TTS, a framework that enables Text‑to‑Speech systems to attribute synthesized speech to their source models without embedding explicit watermarks. By jointly training the TTS model and a discriminator, the method improves traceability generalization while maintaining or slightly enhancing audio quality. This represents the first attempt at watermark‑free TTS with strong traceability, and the authors plan to release the code to support further research.
By Yuxiang Zhao, Yunchong Xiao, Yushen Chen, Zhikang Niu, Shuai Wang, Kai Yu, Xie Chen
CRAW is a codec‑robust audio watermarking framework designed to embed imperceptible signals into synthetic speech, enabling provenance verification. It improves robustness against neural re‑synthesis, codecs, denoisers, and vocoders while preserving perceptual quality through distortion‑aware training, attention‑based pooling, perceptual masking, and error‑correcting codes. Experiments show CRAW outperforms existing post‑hoc watermarking methods in robustness without compromising audio quality.
arXiv:2606. 11828v1 Announce Type: cross Abstract: Audio watermarking aims to embed identifiable information into audio while remaining imperceptible.
By Haiyun Li, Shuhai Peng, Zhisheng Zhang, Jingran Xie, Xiaofeng Xie, Hanyang Peng, Zhiyong Wu
Audio watermarking aims to embed identifiable information into audio while remaining imperceptible. Existing methods adopt high-fidelity, low-energy designs to preserve perceptual quality, but the resulting watermarks lack robustness under suppression by speech reconstruction models.
arXiv:2503.05794v4 Announce Type: replace-cross
Abstract: Speaker verification models are trained on large-scale public datasets whose licenses usually prohibit unauthorized commercial use, yet such...
By Yiming Li, Kaiying Yan, Jiawen Diao, Shuo Shao, Tongqing Zhai, Shu-Tao Xia, Dacheng Tao
arXiv:2603. 14033v2 Announce Type: replace-cross Abstract: Audio anti-spoofing systems are typically trained to assign one authenticity label to an entire speech utterance.
By Shree Harsha Bokkahalli Satish, Harm Lameris, Joakim Gustafson, \'Eva Sz\'ekely
arXiv:2607. 11117v1 Announce Type: cross Abstract: AI music generation has rapidly advanced alongside commercial platforms, raising the need for reliable watermarking for provenance and attribution.
By Seohwan Yun, Jeeyoung Yun, Yongjin Kim, Juyeon Lee, Sungwoong Kim
arXiv:2606. 18430v1 Announce Type: new Abstract: Statistical watermarks help organizations attribute large language model (LLM) outputs, yet existing detectors often struggle when watermark signals are weak, texts are repetitive, or watermarks are edited.
By Chih-Duo Hong, Yen-Pang Chen, Fang Yu
arXiv:2606. 00613v1 Announce Type: cross Abstract: Watermarking should identify language-model output without degrading quality or limiting verification to the model provider.
By Shinwoo Park, Hyejin Park, Hyeseon An, Yo-Sub Han
arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.
By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang