arXiv Machine Learning

Are Two Datasets Close Enough With Statistical Significance? A Kernel Distributional Closeness Testing Approach

arXiv:2507. 12843v3 Announce Type: replace Abstract: Are two distributions close to each other with statistical significance?

arXiv Machine Learning
Jul 28

Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD

arXiv:2607. 24235v1 Announce Type: cross Abstract: Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others.

By Jose Cribeiro-Ramallo, Florian Kalinke, Zolt\'an Szab\'o
arXiv Machine Learning
Jun 29

Efficient and Stable Multi-Dimensional Kolmogorov-Smirnov Distance

arXiv:2504. 11299v2 Announce Type: replace-cross Abstract: We revisit extending the Kolmogorov-Smirnov distance between probability distributions to the multi-dimensional setting, and make new arguments about the proper way to approach this generalization.

By Peter Matthew Jacobs, Foad Namjoo, Jeff M. Phillips
arXiv Machine Learning
Jul 20

Testing Distributions Against Bounded Distinguishers

arXiv:2607. 15645v1 Announce Type: cross Abstract: Motivated by the challenge of testing distributions over high-dimensional or continuous domains, we study distribution testing with respect to bounded classes of distinguishers.

By Mark Bun, Rathin Desai, Renato Ferreira Pinto Jr
arXiv Statistics ML
3d ago

PTED: A multi-dimensional two-sample test for scientific inference and generative machine learning

The article introduces PTED, a Python implementation of a permutation test based on the Energy Distance for two-sample testing in multiple dimensions. PTED uses pairwise distances to compute a test statistic that works in high dimensions, on learned feature representations, and for any data type where a distance can be defined. The authors demonstrate that PTED scales linearly with dimensions and sample size while retaining strong discriminative power, and show it outperforms other multi‑dimensional tests in sensitivity.

By Connor Stone
arXiv Machine Learning
Sep 24

A Discrepancy-Based Perspective on Dataset Condensation

The paper introduces a unified framework for dataset condensation (DC) that generalizes existing methods by using discrepancy measures to quantify the distance between probability distributions. It extends the traditional goal of DC—creating a small synthetic dataset that preserves generalization—to include additional objectives such as robustness and privacy. The framework positions DC as a formal approximation problem, broadening its applicability across different machine learning regimes.

By Tong Chen, Raghavendra Selvan