arXiv Machine Learning

Multi-Channel Spread-Spectrum Code Watermarking

arXiv:2607. 06009v1 Announce Type: cross Abstract: Attributing code to the large language model that produced it is essential for provenance, licensing, and misuse accountability, yet no deployed watermark meets this need.

arXiv AI
Sep 1

Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice

The paper demonstrates that N‑gram based code watermarking schemes, widely used to identify machine‑generated code, are ineffective when faced with realistic code obfuscation. By modeling semantics‑preserving transformations as a Markov random walk and introducing the assumption of distribution consistency, the authors prove that obfuscation can drive the failure rate of any detector to nearly 1 minus its false‑positive rate. Extensive experiments across multiple watermarking methods, LLMs, languages, benchmarks, and obfuscators confirm that detectors collapse to near‑random performance (AUROC ≈ 0.5) after obfuscation.

By Gehao Zhang, Mingzhe Li, Eugene Bagdasarian, Shiqing Ma, Juan Zhai
arXiv Machine Learning
Sep 3

WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading

WeaveMark is a new multi‑bit watermarking scheme for large language models that improves payload capacity, extraction accuracy, and text quality by using coded payload spreading, soft‑decision error‑correcting codes, and unbiased multilayer reweighting. It also adds zero‑bit layers for reliable detection of watermark presence. Experiments demonstrate significant gains, achieving an 89.8% match rate for 32‑bit messages at 200 tokens and maintaining 86.0% accuracy under 10% substitution attacks on 16‑bit messages, far outperforming the BiMark baseline.

By Gang-Hyun Park, Ju-Hyeong Lee, Hee-Youl Kwak, Dae-Young Yun
Hugging Face Trending Papers
Sep 3

Flip, Don't Shuffle: Watermarking LLMs at the Speed of Inference

We introduce Stateless Bernoulli Watermarking (SBW), a new statistical watermark for Large Language Models that determines green list membership through independent per-token Bernoulli trials. Unlike KGW's vocabulary permutation or SynthID's multi-layer tournament, SBW requires only a single comparison per token against a counter-based random number generator, reducing membership complexity to $O(1)$ and enabling single-kernel execution with zero intermediate allocations.