arXiv Machine Learning By Pengrun Huang, Kamalika Chaudhuri, Yu-Xiang Wang

Dataset Watermarking with Provable Black-Box Detection

Read the original on arXiv Machine Learning →

The paper introduces a dataset watermarking technique that embeds a watermark by increasing the co‑occurrence of randomly selected word pairs through meaning‑preserving local edits. The watermark can be detected solely from generated text with provable false‑positive control, and experiments on four base models and three datasets show reliable detection (p < 0.01) even when the watermarked data constitutes less than 5% of fine‑tuning tokens. Compared to existing methods, the approach better preserves benchmark utility and semantic integrity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.