arXiv Machine Learning

Minimax bounds for watermarked and masked recursive discrete distribution estimation

The paper investigates how watermarking affects recursive discrete distribution estimation when synthetic samples are mixed with real data. It establishes minimax lower bounds showing that, as the proportion of real samples approaches zero, adding watermarks cannot improve performance unless the false‑negative detection rate also vanishes. The authors further demonstrate that simple deterministic estimators achieve worst‑case losses close to these bounds and introduce a masking technique that reduces the remaining performance gap to a Jensen gap, suggesting potential for tighter bounds.

arXiv Machine Learning
Jul 7

A Gradient Flow Perspective on Minimum MMD Estimation

arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.

By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen
arXiv Machine Learning
Sep 10

DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory

DRIFT is a black‑box attack that removes diffusion watermarks by deflecting the generative trajectory. It combines partial forward diffusion with stochastic reverse resampling to limit the source information available to a fixed‑depth recovery pipeline and to explore alternative noise‑driven paths. Across nine watermarks, DRIFT achieves 98–100% success while preserving image quality, without requiring secret keys, verifier internals, or per‑image gradient optimization.

By Rui Bao, Zheng Gao, Xiaoyu Li, Xiaoyan Feng, Yang Song, Jiaojiao Jiang
arXiv Machine Learning
Sep 22

On the Information-Theoretic Limits of Latent-Space Watermarking Through Pretrained Generators

The paper investigates latent‑space watermarking using pretrained generators, where a watermark encoder selects latent inputs based on a message and secret key to produce outputs with a specified conditional distribution. For finite alphabets, it derives inner and outer bounds on the rate–key trade‑off and characterizes the capacity region when the generator’s output uniquely determines the latent distribution. The study extends to jointly Gaussian models, identifies key sufficient statistics, optimally allocates secret‑key resources across modes, and analyzes robustness against regeneration attacks, providing compound capacity results and decay rates for repeated attacks.

By Jinwan Jeon, Minju Lee, Sung Hoon Lim
arXiv Machine Learning
Jun 26

Learning from a Biased Sample

arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.

By Roshni Sahoo, Lihua Lei, Stefan Wager
Hugging Face Trending Papers
Sep 8

DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory

DRIFT is a black‑box attack that removes diffusion watermarks by combining partial forward diffusion with stochastic reverse resampling. It limits the source information available to a fixed‑depth recovery pipeline and uses stochastic reversal to explore alternative noise‑driven paths, refining fidelity only on updates rejected by the same verifier. Across nine watermarks, DRIFT achieves 98–100% attack success and the best image quality without requiring secret keys, verifier internals, or per‑image gradient optimization.

arXiv Machine Learning
Jul 10

Prediction-Powered Active Testing

arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.

By Kianoosh Ashouritaklimi, Valentin Kilian, Daolang Huang, Tom Rainforth, Fran\c{c}ois Caron