The paper introduces CertDW, a certified dataset watermark and ownership verification method that remains reliable even under malicious perturbations. By leveraging conformal prediction, it defines two statistical measures—principal probability (PP) and watermark robustness (WR)—to evaluate model stability on benign versus watermarked samples. The authors derive certification conditions linking WR to a PP-based threshold and provide a high‑probability bound on false positives, enabling robust ownership verification when a suspicious model’s WR exceeds the PP values of benign models.
By Ting Qiao, Yiming Li, Jianbin Li, Yingjia Wang, Leyi Qi, Junfeng Guo, Ruili Feng, Dacheng Tao
arXiv:2606. 23335v2 Announce Type: replace-cross Abstract: Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models such as Chatterbox, provided through dedicated techniques such as AudioSeal, or deployed by commercial platforms such as ElevenLabs.
By Nicolas M. M\"uller, Pascal Debus
Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models such as Chatterbox, provided through dedicated techniques such as AudioSeal, or deployed by commercial platforms such as ElevenLabs. We identify a previously uncharacterized liability: when synthetic speech is watermarked and human speech is not, detectors trained alongside latch onto the watermark as a spurious "watermark => fake" shortcut.
arXiv:2501. 15509v5 Announce Type: replace-cross Abstract: Model fingerprinting has emerged as a crucial mechanism for safeguarding the intellectual property of open-source models, offering a non-intrusive approach that requires no modifications to the protected model.
By Shuo Shao, Haozhe Zhu, Yiming Li, Hongwei Yao, Tianwei Zhang, Zhan Qin
arXiv:2606. 18430v1 Announce Type: new Abstract: Statistical watermarks help organizations attribute large language model (LLM) outputs, yet existing detectors often struggle when watermark signals are weak, texts are repetitive, or watermarks are edited.
By Chih-Duo Hong, Yen-Pang Chen, Fang Yu
arXiv:2603. 23171v3 Announce Type: replace-cross Abstract: Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent.
By Toluwani Aremu, Daniil Ognev, Samuele Poppi, Nils Lukas
arXiv:2606. 11698v1 Announce Type: cross Abstract: Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures.
By Jian-Ping Mei, Weibin Zhang, Ao Yao, Tiantian Zhu, Jie Xiao
arXiv:2607. 00325v1 Announce Type: new Abstract: A growing body of literature suggests that training data membership inference problems are fundamentally hard tasks in modern language modeling settings.
By John Kirchenbauer, Brian R. Bartoldson, Bhavya Kailkhura, Tom Goldstein
The paper introduces REMARK, a watermark‑based fingerprint framework designed to verify ownership of Graph Neural Networks (GNNs). REMARK generates in‑distribution watermark graphs that maximize output differences between GNN models, thereby reducing performance loss from out‑of‑distribution watermarks. It then extracts robust fingerprints from these output differences, eliminating the need for surrogate models trained on watermark data or reliance on specific output levels, and achieves state‑of‑the‑art verification accuracy across real‑world datasets and GNN architectures.
By Han Zhang, Yan Wang, Guanfeng Liu, Pengfei Ding, Huaxiong Wang, Kwok-Yan Lam
arXiv:2607. 22035v1 Announce Type: new Abstract: Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain concepts, common stylistic conventions, or ordinary statistical generalization.
By Xiafeng Man
arXiv:2510. 10982v2 Announce Type: replace-cross Abstract: Recent AI regulations increasingly emphasize the need for mechanisms that preserve the utility of data for AI innovation while preventing misuse, particularly by enforcing purpose limitation in downstream AI applications.
By Zihan Wang, Zhiyong Ma, Zhongkui Ma, Shuofeng Liu, Akide Liu, Derui Wang, Minhui Xue, Guangdong Bai
arXiv:2605. 07663v2 Announce Type: replace-cross Abstract: Data valuation methods allocate payments and audit training data's contribution to machine-learning pipelines; however, they often assume passive contributors.
By Florian A. D. Burnat, Brittany I. Davidson