arXiv Computation and Language

More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles

The paper investigates watermarking for large language model outputs, noting that stronger single-layer watermarks reduce token entropy and weaken subsequent layers. It demonstrates that detectability is limited by entropy and that watermark ensembles monotonically lower entropy and the green‑list ratio. The authors propose using weaker single‑layer watermarks to maintain entropy, showing through theory and experiments that this approach improves both detectability and robustness compared to strong baselines.

arXiv Machine Learning
Jul 8

Beyond Heuristic Tuning: Power-Calibrated LLM Watermarking

arXiv:2607. 05694v1 Announce Type: cross Abstract: Logit-based watermarking is a widely used mechanism for identifying LLM generated content, yet its effectiveness is governed by a fundamental trade-off between detectability and semantic distortion.

By Xiaopu Wang, Zelin He, Chengyuan Liu, Runze Li
arXiv Computation and Language
4d ago

Semantic Watermarking with Order-Robust Detection over Sub-sentence Units

The paper introduces an adaptive embedding displacement attack (EDA) that exploits rewording, reordering, and resegmentation to remove semantic watermarks from text, achieving a 32.6%–47.9% success rate across four watermarking schemes. To counter this, the authors propose k‑SwordStamp, a semantic watermarking method that uses order‑robust detection over sub‑sentence units, significantly reducing vulnerability to structure‑based edits. Experiments show that EDA remains effective against k‑SwordStamp, but with a lower success rate (10.8%) compared to its performance on other schemes.

By Abdulrahman Diaa, Jonathan Petit, Florian Kerschbaum