arXiv Machine Learning

Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics

arXiv:2512. 13997v2 Announce Type: replace-cross Abstract: Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions.

arXiv Machine Learning
Sep 23

Conditional Distributional Treatment Effects: Doubly Robust Estimation and Testing

The paper introduces a new estimand for conditional distributional treatment effects that captures how treatments influence the entire outcome distribution, including variance and tail risks, in a covariate-dependent manner. It presents a doubly robust estimator that is minimax optimal locally and uses it to construct a test for global homogeneity of conditional potential outcome distributions. The test accommodates discrepancies beyond the maximum mean discrepancy, guarantees valid type‑1 error, is consistent against fixed alternatives, and includes a computationally efficient, permutation‑free algorithm with exact closed‑form expressions for two natural discrepancies.

By Saksham Jain, Alex Luedtke
arXiv Machine Learning
Jul 28

Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD

arXiv:2607. 24235v1 Announce Type: cross Abstract: Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others.

By Jose Cribeiro-Ramallo, Florian Kalinke, Zolt\'an Szab\'o
arXiv Statistics ML
Sep 7

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models

The paper investigates Double Machine Learning (DML) estimators under structure‑agnostic (SA) models, which assume the data‑generating law lies within a neighborhood of fixed machine‑learning estimates. It shows that for two of three studied functionals—the quadratic functional in the Gaussian sequence model and the quadratic density integral functional—the DML estimators are asymptotically inadmissible, being dominated by second‑order empirical higher‑order influence function (HOIF) estimators. For the third functional, the expected conditional covariance, both DML and HOIF estimators remain minimax but neither dominates the other.

By Lin Liu, Rajarshi Mukherjee, James M Robins
arXiv Statistics ML
4d ago

Bentkus-type asymptotic e-values

arXiv:2606.06332v2 Announce Type: replace-cross Abstract: Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, w...

By Diego Martinez-Taboada, Ben Chugg, Aaditya Ramdas
arXiv Machine Learning
Sep 1

Minimax bounds for watermarked and masked recursive discrete distribution estimation

The paper investigates how watermarking affects recursive discrete distribution estimation when synthetic samples are mixed with real data. It establishes minimax lower bounds showing that, as the proportion of real samples approaches zero, adding watermarks cannot improve performance unless the false‑negative detection rate also vanishes. The authors further demonstrate that simple deterministic estimators achieve worst‑case losses close to these bounds and introduce a masking technique that reduces the remaining performance gap to a Jensen gap, suggesting potential for tighter bounds.

By Millen Kanabar, Michael Gastpar