arXiv AI By Xuyang Chen, Xiang Li, Yangxinyu Xie, Qi Long

Selective Disclosure Watermarking for Large Language Models

Read the original on arXiv AI →

arXiv:2607. 05353v1 Announce Type: cross Abstract: Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading

WeaveMark is a new multi‑bit watermarking scheme for large language models that improves payload capacity, extraction accuracy, and text quality by using coded payload spreading, soft‑decision error‑correcting codes, and unbiased multilayer reweighting. It also adds zero‑bit layers for reliable detection of watermark presence. Experiments demonstrate significant gains, achieving an 89.8% match rate for 32‑bit messages at 200 tokens and maintaining 86.0% accuracy under 10% substitution attacks on 16‑bit messages, far outperforming the BiMark baseline.

By Gang-Hyun Park, Ju-Hyeong Lee, Hee-Youl Kwak, Dae-Young Yun
arXiv Computation and Language
Sep 7

Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings

The paper introduces Dual-Embedding Watermarking (DEW), a semantic watermarking technique for large language models that combines contextual and token-level embeddings. DEW applies algebraic vector-space operations to generate a watermark signal that remains robust to paraphrasing and translation, while obfuscating the signal with pseudo-random matrices seeded by a secret key. Experiments demonstrate state‑of‑the‑art robustness, especially against translation, with minimal computational overhead and preserved text quality at lower watermark strengths.

By Jonas Sch\"afer, Cezary Pilaszewicz, Gerhard Wunder