arXiv Machine Learning

Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS

arXiv Statistics ML
3d ago

PTED: A multi-dimensional two-sample test for scientific inference and generative machine learning

The article introduces PTED, a Python implementation of a permutation test based on the Energy Distance for two-sample testing in multiple dimensions. PTED uses pairwise distances to compute a test statistic that works in high dimensions, on learned feature representations, and for any data type where a distance can be defined. The authors demonstrate that PTED scales linearly with dimensions and sample size while retaining strong discriminative power, and show it outperforms other multi‑dimensional tests in sensitivity.

By Connor Stone
arXiv Machine Learning
Aug 19

Feature Priming in Online Linear Regression: Sparse-Regret Lower Bounds and a Tight Univariate Rate

The paper investigates feature priming in high‑dimensional online linear regression, showing that estimating feature weights from past data and refitting a minimum‑norm predictor can lead to regret that scales with sparsity rather than ambient dimension. It provides a negative answer to a COLT 2023 open problem by proving that three natural priming rules incur ≥Ω(min{T,√d}) regret against a zero‑loss one‑sparse comparator, due to cheap nuisance interpolation that underweights truly predictive coordinates. The authors also identify conditions under which regret is governed by data rank and present constructions that achieve tight univariate rates, while noting that the multivariate case remains unresolved.

By Huibo Xu, Shi Fu, Qixin Zhang, Dacheng Tao
arXiv Machine Learning
Jun 9

Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models

arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.

By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong