arXiv AI

Beyond Semantic Narrowing: Robust and Efficient LLM Watermarking with Hamming Neighborhoods

arXiv Machine Learning
2d ago

Dataset Watermarking with Provable Black-Box Detection

The paper introduces a dataset watermarking technique that embeds a watermark by increasing the co‑occurrence of randomly selected word pairs through meaning‑preserving local edits. The watermark can be detected solely from generated text with provable false‑positive control, and experiments on four base models and three datasets show reliable detection (p < 0.01) even when the watermarked data constitutes less than 5% of fine‑tuning tokens. Compared to existing methods, the approach better preserves benchmark utility and semantic integrity.

By Pengrun Huang, Kamalika Chaudhuri, Yu-Xiang Wang
arXiv Computation and Language
Aug 31

Semantic Watermarking with Order-Robust Detection over Sub-sentence Units

The paper introduces an adaptive embedding displacement attack (EDA) that exploits rewording, reordering, and resegmentation to remove semantic watermarks from text, achieving a 32.6%–47.9% success rate across four watermarking schemes. To counter this, the authors propose k‑SwordStamp, a semantic watermarking method that uses order‑robust detection over sub‑sentence units, significantly reducing vulnerability to structure‑based edits. Experiments show that EDA remains effective against k‑SwordStamp, but with a lower success rate (10.8%) compared to its performance on other schemes.

By Abdulrahman Diaa, Jonathan Petit, Florian Kerschbaum
arXiv Computation and Language
2d ago

TTMark: Pairwise Distortion-Free Watermarking Beyond Single-Token Entropy

TTMark introduces a pairwise watermarking framework that extends distortion‑free watermarking from single tokens to adjacent token pairs, enlarging the watermarking alphabet from V to V². By watermarking the joint distribution of consecutive tokens, the detector can exploit both token entropy and conditional entropy while maintaining distortion‑freeness. Experiments on multiple language models and datasets show that TTMARK improves detectability, robustness to edits, and localized watermark detection without degrading generation quality.

By Ruibo Chen, Zhengmian Hu, Donghang Lu, Xuehao Cui, Georgios Milis, Yihan Wu, Jian Du, Heng Huang
arXiv Computation and Language
Sep 7

Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings

The paper introduces Dual-Embedding Watermarking (DEW), a semantic watermarking technique for large language models that combines contextual and token-level embeddings. DEW applies algebraic vector-space operations to generate a watermark signal that remains robust to paraphrasing and translation, while obfuscating the signal with pseudo-random matrices seeded by a secret key. Experiments demonstrate state‑of‑the‑art robustness, especially against translation, with minimal computational overhead and preserved text quality at lower watermark strengths.

By Jonas Sch\"afer, Cezary Pilaszewicz, Gerhard Wunder
arXiv Computation and Language
2d ago

CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking

CORE-BREW is a new multi‑bit watermarking method for large language models that uses log‑likelihood ratios for soft‑decision decoding, targeting a fixed hit rate to calibrate the watermark channel. It introduces entropy‑aware erasures to reduce perturbations in low‑entropy contexts and combines likelihood‑based scoring with soft‑decision list decoding to better exploit token‑level reliability. Experiments on open‑source LLMs show that CORE‑BREW improves detection robustness and payload recovery compared to the BREW baseline while keeping false‑positive rates low and maintaining translation quality metrics close to unwatermarked text.

By Joeun Kim, HoEun Kim, Young-Sik Kim