arXiv AI

Selective Disclosure Watermarking for Large Language Models

arXiv:2607. 05353v1 Announce Type: cross Abstract: Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs).

arXiv Machine Learning
Sep 3

WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading

WeaveMark is a new multi‑bit watermarking scheme for large language models that improves payload capacity, extraction accuracy, and text quality by using coded payload spreading, soft‑decision error‑correcting codes, and unbiased multilayer reweighting. It also adds zero‑bit layers for reliable detection of watermark presence. Experiments demonstrate significant gains, achieving an 89.8% match rate for 32‑bit messages at 200 tokens and maintaining 86.0% accuracy under 10% substitution attacks on 16‑bit messages, far outperforming the BiMark baseline.

By Gang-Hyun Park, Ju-Hyeong Lee, Hee-Youl Kwak, Dae-Young Yun
arXiv Computation and Language
Sep 7

Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings

The paper introduces Dual-Embedding Watermarking (DEW), a semantic watermarking technique for large language models that combines contextual and token-level embeddings. DEW applies algebraic vector-space operations to generate a watermark signal that remains robust to paraphrasing and translation, while obfuscating the signal with pseudo-random matrices seeded by a secret key. Experiments demonstrate state‑of‑the‑art robustness, especially against translation, with minimal computational overhead and preserved text quality at lower watermark strengths.

By Jonas Sch\"afer, Cezary Pilaszewicz, Gerhard Wunder
arXiv Machine Learning
Jul 8

Multi-Channel Spread-Spectrum Code Watermarking

arXiv:2607. 06009v1 Announce Type: cross Abstract: Attributing code to the large language model that produced it is essential for provenance, licensing, and misuse accountability, yet no deployed watermark meets this need.

By Soohyeon Choi, Debin Gao, Yue Duan
arXiv Machine Learning
4d ago

Dataset Watermarking with Provable Black-Box Detection

The paper introduces a dataset watermarking technique that embeds a watermark by increasing the co‑occurrence of randomly selected word pairs through meaning‑preserving local edits. The watermark can be detected solely from generated text with provable false‑positive control, and experiments on four base models and three datasets show reliable detection (p < 0.01) even when the watermarked data constitutes less than 5% of fine‑tuning tokens. Compared to existing methods, the approach better preserves benchmark utility and semantic integrity.

By Pengrun Huang, Kamalika Chaudhuri, Yu-Xiang Wang
arXiv AI
Sep 1

Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice

The paper demonstrates that N‑gram based code watermarking schemes, widely used to identify machine‑generated code, are ineffective when faced with realistic code obfuscation. By modeling semantics‑preserving transformations as a Markov random walk and introducing the assumption of distribution consistency, the authors prove that obfuscation can drive the failure rate of any detector to nearly 1 minus its false‑positive rate. Extensive experiments across multiple watermarking methods, LLMs, languages, benchmarks, and obfuscators confirm that detectors collapse to near‑random performance (AUROC ≈ 0.5) after obfuscation.

By Gehao Zhang, Mingzhe Li, Eugene Bagdasarian, Shiqing Ma, Juan Zhai