arXiv Machine Learning

Sample Complexities of Estimating Gumbel--Max Watermark Proportions with and without Reduction to Pivotal Statistics

arXiv:2607. 00224v1 Announce Type: cross Abstract: Watermarking promises a statistical trace of large language model (LLM) use, but real documents, after editing or paraphrasing, rarely arrive as purely human-written or purely machine-generated.

arXiv AI
Sep 18

Watermarking Diffusion Language Models

The paper introduces the first watermark designed specifically for diffusion language models (DLMs), which generate tokens in arbitrary order unlike traditional autoregressive models. It overcomes the challenge of missing prior tokens by applying the watermark in expectation over the context and promoting tokens that strengthen the watermark when used as context. Experiments show a >99% true positive rate with minimal quality loss and comparable robustness to existing autoregressive watermarks.

By Thibaud Gloaguen, Robin Staab, Nikola Jovanovi\'c, Martin Vechev
arXiv Machine Learning
Jul 8

Beyond Heuristic Tuning: Power-Calibrated LLM Watermarking

arXiv:2607. 05694v1 Announce Type: cross Abstract: Logit-based watermarking is a widely used mechanism for identifying LLM generated content, yet its effectiveness is governed by a fundamental trade-off between detectability and semantic distortion.

By Xiaopu Wang, Zelin He, Chengyuan Liu, Runze Li
arXiv Machine Learning
Sep 1

Minimax bounds for watermarked and masked recursive discrete distribution estimation

The paper investigates how watermarking affects recursive discrete distribution estimation when synthetic samples are mixed with real data. It establishes minimax lower bounds showing that, as the proportion of real samples approaches zero, adding watermarks cannot improve performance unless the false‑negative detection rate also vanishes. The authors further demonstrate that simple deterministic estimators achieve worst‑case losses close to these bounds and introduce a masking technique that reduces the remaining performance gap to a Jensen gap, suggesting potential for tighter bounds.

By Millen Kanabar, Michael Gastpar
arXiv Machine Learning
Sep 21

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

The paper introduces a new multi-draft speculative sampling algorithm that uses Poisson processes to improve inference efficiency and output provenance for large language models. It achieves strong sampling efficiency while allowing an unbiased watermark to be embedded without reducing speculative acceptance. The method relies on an exact list‑coupling‑without‑communication scheme, giving a drafter‑invariant property that benefits both sampling and watermarking, and the authors experimentally confirm its effectiveness.

By Yanxiao Liu, Sicheng Wan, Zhan Gao, Deniz G\"und\"uz