arXiv:2601. 11629v2 Announce Type: replace-cross Abstract: We demonstrate that while the current approaches for language model watermarking are effective for open-ended generation, they are inadequate at watermarking LM outputs for constrained generation tasks with low-entropy output spaces.
By Nghia T. Le, Alan Ritter, Kartik Goyal
The paper introduces Dual-Embedding Watermarking (DEW), a semantic watermarking technique for large language models that combines contextual and token-level embeddings. DEW applies algebraic vector-space operations to generate a watermark signal that remains robust to paraphrasing and translation, while obfuscating the signal with pseudo-random matrices seeded by a secret key. Experiments demonstrate state‑of‑the‑art robustness, especially against translation, with minimal computational overhead and preserved text quality at lower watermark strengths.
By Jonas Sch\"afer, Cezary Pilaszewicz, Gerhard Wunder
WeaveMark is a new multi‑bit watermarking scheme for large language models that improves payload capacity, extraction accuracy, and text quality by using coded payload spreading, soft‑decision error‑correcting codes, and unbiased multilayer reweighting. It also adds zero‑bit layers for reliable detection of watermark presence. Experiments demonstrate significant gains, achieving an 89.8% match rate for 32‑bit messages at 200 tokens and maintaining 86.0% accuracy under 10% substitution attacks on 16‑bit messages, far outperforming the BiMark baseline.
By Gang-Hyun Park, Ju-Hyeong Lee, Hee-Youl Kwak, Dae-Young Yun
The paper introduces the first watermark designed specifically for diffusion language models (DLMs), which generate tokens in arbitrary order unlike traditional autoregressive models. It overcomes the challenge of missing prior tokens by applying the watermark in expectation over the context and promoting tokens that strengthen the watermark when used as context. Experiments show a >99% true positive rate with minimal quality loss and comparable robustness to existing autoregressive watermarks.
By Thibaud Gloaguen, Robin Staab, Nikola Jovanovi\'c, Martin Vechev
The paper investigates watermarking for large language model outputs, noting that stronger single-layer watermarks reduce token entropy and weaken subsequent layers. It demonstrates that detectability is limited by entropy and that watermark ensembles monotonically lower entropy and the green‑list ratio. The authors propose using weaker single‑layer watermarks to maintain entropy, showing through theory and experiments that this approach improves both detectability and robustness compared to strong baselines.
By Ruibo Chen, Yihan Wu, Xuehao Cui, Jingqi Zhang, Heng Huang
CORE-BREW is a new multi‑bit watermarking method for large language models that uses log‑likelihood ratios for soft‑decision decoding, targeting a fixed hit rate to calibrate the watermark channel. It introduces entropy‑aware erasures to reduce perturbations in low‑entropy contexts and combines likelihood‑based scoring with soft‑decision list decoding to better exploit token‑level reliability. Experiments on open‑source LLMs show that CORE‑BREW improves detection robustness and payload recovery compared to the BREW baseline while keeping false‑positive rates low and maintaining translation quality metrics close to unwatermarked text.
By Joeun Kim, HoEun Kim, Young-Sik Kim