arXiv AI

WorldMark: A Plug-and-Play World Knowledge Interface for Cross-Host Language Model Watermarking

arXiv:2608. 06416v1 Announce Type: cross Abstract: Watermarking traces the provenance of text produced by large language models by embedding statistically detectable signals during decoding.

arXiv Computation and Language
4d ago

CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking

CORE-BREW is a new multi‑bit watermarking method for large language models that uses log‑likelihood ratios for soft‑decision decoding, targeting a fixed hit rate to calibrate the watermark channel. It introduces entropy‑aware erasures to reduce perturbations in low‑entropy contexts and combines likelihood‑based scoring with soft‑decision list decoding to better exploit token‑level reliability. Experiments on open‑source LLMs show that CORE‑BREW improves detection robustness and payload recovery compared to the BREW baseline while keeping false‑positive rates low and maintaining translation quality metrics close to unwatermarked text.

By Joeun Kim, HoEun Kim, Young-Sik Kim
arXiv Machine Learning
4d ago

Dataset Watermarking with Provable Black-Box Detection

The paper introduces a dataset watermarking technique that embeds a watermark by increasing the co‑occurrence of randomly selected word pairs through meaning‑preserving local edits. The watermark can be detected solely from generated text with provable false‑positive control, and experiments on four base models and three datasets show reliable detection (p < 0.01) even when the watermarked data constitutes less than 5% of fine‑tuning tokens. Compared to existing methods, the approach better preserves benchmark utility and semantic integrity.

By Pengrun Huang, Kamalika Chaudhuri, Yu-Xiang Wang
arXiv Computation and Language
Sep 1

More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles

The paper investigates watermarking for large language model outputs, noting that stronger single-layer watermarks reduce token entropy and weaken subsequent layers. It demonstrates that detectability is limited by entropy and that watermark ensembles monotonically lower entropy and the green‑list ratio. The authors propose using weaker single‑layer watermarks to maintain entropy, showing through theory and experiments that this approach improves both detectability and robustness compared to strong baselines.

By Ruibo Chen, Yihan Wu, Xuehao Cui, Jingqi Zhang, Heng Huang
arXiv Machine Learning
Aug 20

Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text

The paper introduces Pattern Stability Score (PSS), a watermark detection framework that uses local statistical features and stability dynamics across paraphrased variants to identify machine-generated text. PSS combines global and local z‑score features with higher‑order run‑length statistics, autocorrelation signals, and stability scores over paraphrase depth. Experiments on PG‑19, CNN/DailyMail, and WikiText with Llama‑3‑8B, Qwen2‑7B, and multiple paraphrasers show that PSS improves detection AUC by 10‑15 percentage points and a single universal classifier achieves over 87.8% AUC across diverse LLMs, paraphrasers, and domains without retraining.

By Sina Mansouri, Mohit Marvania, Abolfazl Safikhani