arXiv:2609.36307v1 Announce Type: new
Abstract: Partial Least Squares (PLS) regression extracts a few outcome-aligned directions in a high-dimensional X and is widely used across applied science, but...
By Pawe{\l} Lenartowicz, Hubert Plisiecki
arXiv:2607. 15645v1 Announce Type: cross Abstract: Motivated by the challenge of testing distributions over high-dimensional or continuous domains, we study distribution testing with respect to bounded classes of distinguishers.
By Mark Bun, Rathin Desai, Renato Ferreira Pinto Jr
arXiv:2604. 05324v2 Announce Type: replace Abstract: Statistical evaluation aims to estimate the generalization performance of a model using held-out i.
By Shashaank Aiyer, Yishay Mansour, Shay Moran, Han Shao
arXiv:2507. 12843v3 Announce Type: replace Abstract: Are two distributions close to each other with statistical significance?
By Zhijian Zhou, Liuhua Peng, Xunye Tian, Mingming Gong, Feng Liu
arXiv:2405. 07780v3 Announce Type: replace-cross Abstract: This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and arbitrarily imbalanced.
By Zhiyong Yang, Qianqian Xu, Sicong Li, Zitai Wang, Xiaochun Cao, Qingming Huang
arXiv:2504. 11299v2 Announce Type: replace-cross Abstract: We revisit extending the Kolmogorov-Smirnov distance between probability distributions to the multi-dimensional setting, and make new arguments about the proper way to approach this generalization.
By Peter Matthew Jacobs, Foad Namjoo, Jeff M. Phillips
arXiv:2606. 04009v1 Announce Type: cross Abstract: Two-sample testing is a fundamental tool for detecting distributional differences across scientific domains, but classical tests (including kernel-based tests) can be ineffective on high-dimensional structured data such as images.
By Wei-Cheng Lai, Marco Simnacher, Christoph Lippert
arXiv:2512. 13997v2 Announce Type: replace-cross Abstract: Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions.
By Aaron Wei, Milad Jalali, Danica J. Sutherland
The paper critiques current memorization audits for generative models, arguing that lacking a proper null distribution leads to misleading conclusions. It introduces two exact null tests—one permutation test for whole models and a calibrated test for single images—showing that many previously flagged memorizations disappear under these stricter controls. The authors also propose a scale‑restricted statistic based on the Intersection Euler Characteristic Profile to better detect distinct copied images.
By Sushovan Majhi, Pramita Bagchi
The paper introduces View distance, a novel metric that projects high‑dimensional data onto all pairwise two‑dimensional planes and sums the Euclidean distances across these projections. It satisfies metric axioms, couples features, suppresses redundancy, and captures anisotropic geometry. To make it scalable, the authors propose a plane‑selection strategy using iterative Maximum Weight Matching, reducing complexity from ω(n²) to ω(k) and demonstrating competitive performance on twelve datasets.
By Yiqun Zhang, Hou-biao Li
arXiv:2606. 24178v1 Announce Type: cross Abstract: Pretrained vision models often misclassify inputs that are rotated, scaled, or sheared, even though these affine transformations leave the object class unchanged.
By Dominik Lindner, Johann Schmidt, Tom Siegl, Martin Becker, Sebastian Stober
arXiv:2608.24881v1 Announce Type: cross
Abstract: Generative models are commonly ranked by Fr\'echet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary c...
By Hao Chen