arXiv:2607. 00224v1 Announce Type: cross Abstract: Watermarking promises a statistical trace of large language model (LLM) use, but real documents, after editing or paraphrasing, rarely arrive as purely human-written or purely machine-generated.
By Shuwen Chai, Qiaosen Wang
arXiv:2607. 22889v1 Announce Type: new Abstract: Learning the natural parameters $z \in \mathbb{R}^n$ of discrete distributions $\mu_z$ from independent samples constrained to a subset $S \subseteq \{0,1\}^n$ is a foundational challenge in high-dimensional statistics.
By Rohan Chauhan, Ioannis Panageas
arXiv:2505. 20178v2 Announce Type: replace-cross Abstract: Prediction-Powered Inference (PPI) is a popular strategy for combining gold-standard and possibly noisy pseudo-labels to perform statistical estimation.
By Pranav Mani, Peng Xu, Zachary C. Lipton, Michael Oberst
arXiv:2606. 27685v1 Announce Type: cross Abstract: Pervasive data contamination -- stemming from measurement errors, outliers, or adversarial corruption -- has motivated the development of robust statistical methods.
By Shixiang Liu, Hanming Yang
arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.
By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen
DRIFT is a black‑box attack that removes diffusion watermarks by deflecting the generative trajectory. It combines partial forward diffusion with stochastic reverse resampling to limit the source information available to a fixed‑depth recovery pipeline and to explore alternative noise‑driven paths. Across nine watermarks, DRIFT achieves 98–100% success while preserving image quality, without requiring secret keys, verifier internals, or per‑image gradient optimization.
By Rui Bao, Zheng Gao, Xiaoyu Li, Xiaoyan Feng, Yang Song, Jiaojiao Jiang
The paper investigates latent‑space watermarking using pretrained generators, where a watermark encoder selects latent inputs based on a message and secret key to produce outputs with a specified conditional distribution. For finite alphabets, it derives inner and outer bounds on the rate–key trade‑off and characterizes the capacity region when the generator’s output uniquely determines the latent distribution. The study extends to jointly Gaussian models, identifies key sufficient statistics, optimally allocates secret‑key resources across modes, and analyzes robustness against regeneration attacks, providing compound capacity results and decay rates for repeated attacks.
By Jinwan Jeon, Minju Lee, Sung Hoon Lim
arXiv:2512. 13997v2 Announce Type: replace-cross Abstract: Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions.
By Aaron Wei, Milad Jalali, Danica J. Sutherland
arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.
By Roshni Sahoo, Lihua Lei, Stefan Wager
arXiv:2602. 17894v2 Announce Type: replace-cross Abstract: Data collection is a critical component of modern statistical and machine learning pipelines, particularly when data must be gathered from multiple heterogeneous sources to study a target population of interest.
By Michael O. Harding, Vikas Singh, Kirthevasan Kandasamy
DRIFT is a black‑box attack that removes diffusion watermarks by combining partial forward diffusion with stochastic reverse resampling. It limits the source information available to a fixed‑depth recovery pipeline and uses stochastic reversal to explore alternative noise‑driven paths, refining fidelity only on updates rejected by the same verifier. Across nine watermarks, DRIFT achieves 98–100% attack success and the best image quality without requiring secret keys, verifier internals, or per‑image gradient optimization.
arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.
By Kianoosh Ashouritaklimi, Valentin Kilian, Daolang Huang, Tom Rainforth, Fran\c{c}ois Caron