arXiv Machine Learning

Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes

arXiv:2607. 21958v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed, reliable and efficient mechanisms for distinguishing AI-generated text from human-written content have become essential.

arXiv AI
Sep 18

Watermarking Diffusion Language Models

The paper introduces the first watermark designed specifically for diffusion language models (DLMs), which generate tokens in arbitrary order unlike traditional autoregressive models. It overcomes the challenge of missing prior tokens by applying the watermark in expectation over the context and promoting tokens that strengthen the watermark when used as context. Experiments show a >99% true positive rate with minimal quality loss and comparable robustness to existing autoregressive watermarks.

By Thibaud Gloaguen, Robin Staab, Nikola Jovanovi\'c, Martin Vechev
arXiv Computation and Language
4d ago

CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking

CORE-BREW is a new multi‑bit watermarking method for large language models that uses log‑likelihood ratios for soft‑decision decoding, targeting a fixed hit rate to calibrate the watermark channel. It introduces entropy‑aware erasures to reduce perturbations in low‑entropy contexts and combines likelihood‑based scoring with soft‑decision list decoding to better exploit token‑level reliability. Experiments on open‑source LLMs show that CORE‑BREW improves detection robustness and payload recovery compared to the BREW baseline while keeping false‑positive rates low and maintaining translation quality metrics close to unwatermarked text.

By Joeun Kim, HoEun Kim, Young-Sik Kim
arXiv Machine Learning
Sep 21

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

The paper introduces a new multi-draft speculative sampling algorithm that uses Poisson processes to improve inference efficiency and output provenance for large language models. It achieves strong sampling efficiency while allowing an unbiased watermark to be embedded without reducing speculative acceptance. The method relies on an exact list‑coupling‑without‑communication scheme, giving a drafter‑invariant property that benefits both sampling and watermarking, and the authors experimentally confirm its effectiveness.

By Yanxiao Liu, Sicheng Wan, Zhan Gao, Deniz G\"und\"uz