arXiv Computation and Language

SemTrace: Source-Grounded Semantic Signatures for Tracing LLM Exposure to Protected Documents

arXiv Machine Learning
4d ago

Dataset Watermarking with Provable Black-Box Detection

The paper introduces a dataset watermarking technique that embeds a watermark by increasing the co‑occurrence of randomly selected word pairs through meaning‑preserving local edits. The watermark can be detected solely from generated text with provable false‑positive control, and experiments on four base models and three datasets show reliable detection (p < 0.01) even when the watermarked data constitutes less than 5% of fine‑tuning tokens. Compared to existing methods, the approach better preserves benchmark utility and semantic integrity.

By Pengrun Huang, Kamalika Chaudhuri, Yu-Xiang Wang
arXiv Computation and Language
4d ago

CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking

CORE-BREW is a new multi‑bit watermarking method for large language models that uses log‑likelihood ratios for soft‑decision decoding, targeting a fixed hit rate to calibrate the watermark channel. It introduces entropy‑aware erasures to reduce perturbations in low‑entropy contexts and combines likelihood‑based scoring with soft‑decision list decoding to better exploit token‑level reliability. Experiments on open‑source LLMs show that CORE‑BREW improves detection robustness and payload recovery compared to the BREW baseline while keeping false‑positive rates low and maintaining translation quality metrics close to unwatermarked text.

By Joeun Kim, HoEun Kim, Young-Sik Kim
arXiv Computation and Language
Sep 4

Rent-a-RAG: Embedding-Space Watermarks for Auditing Third-Party RAG

The paper introduces DirBucket, a provider-side semantic watermarking system designed to audit document reuse in third‑party retrieval‑augmented generation (RAG) marketplaces. By embedding watermarks as meaning‑preserving paraphrases biased toward provider‑specific directions, DirBucket enables black‑box detection of unauthorized document reuse while maintaining retrieval quality. Experiments on mixed‑provider benchmarks show DirBucket consistently detects non‑compliance within 23 answers, resists adversarial laundering, and transfers effectively to real‑world corpora.

By Alexandr Goultiaev Tolstokorov, Kyriakos Mouratidis, Javad Dogani, Nikolaos Laoutaris
arXiv AI
Sep 10

AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents

AtomCite is an agentic framework that verifies and corrects page‑level citations in multi‑page documents by parsing answers into claims, checking each claim against the cited page image, and applying a deterministic repair policy. The authors introduce DocCite, the first benchmark for this task, built on MP‑DocVQA and DUDE, containing 928 injected instances and 1,909 verified natural errors. Across Gemini, Claude, and GPT models, AtomCite achieves about 93% verification accuracy and improves citation precision from 34% to 87‑90%, while also enhancing hallucination detection in open‑source models.

By Chen Qian, Yimeng Wang, Yu Chen, Lingfei Wu, Andreas Stathopoulos
arXiv Computation and Language
Aug 31

Semantic Watermarking with Order-Robust Detection over Sub-sentence Units

The paper introduces an adaptive embedding displacement attack (EDA) that exploits rewording, reordering, and resegmentation to remove semantic watermarks from text, achieving a 32.6%–47.9% success rate across four watermarking schemes. To counter this, the authors propose k‑SwordStamp, a semantic watermarking method that uses order‑robust detection over sub‑sentence units, significantly reducing vulnerability to structure‑based edits. Experiments show that EDA remains effective against k‑SwordStamp, but with a lower success rate (10.8%) compared to its performance on other schemes.

By Abdulrahman Diaa, Jonathan Petit, Florian Kerschbaum
Hugging Face Trending Papers
Sep 3

Rent-a-RAG: Embedding-Space Watermarks for Auditing Third-Party RAG

Rent‑a‑RAG introduces DirBucket, a provider‑side semantic watermarking system that embeds documents with meaning‑preserving paraphrases biased toward secret directions, allowing black‑box auditing of document reuse in multi‑provider retrieval‑augmented generation (RAG). The method consistently detects non‑compliance on a challenging benchmark, achieving strong target detection with no false positives and surviving adversarial post‑answer laundering. DirBucket’s detection transfers unchanged to real‑world corpora in clinical, cyber‑threat‑intelligence, and legal domains, demonstrating that embedding‑space watermarking can make third‑party RAG document reuse statistically auditable.