arXiv Statistics ML

Synthetic Nearest Neighbors: Extending Synthetic Controls for Matrix Completion with Missing Not at Random Data

arXiv Machine Learning
Aug 18

Doubly robust nearest neighbors in factor models

arXiv:2211. 14297v4 Announce Type: replace-cross Abstract: We introduce and analyze an improved variant of nearest neighbors (NN) for estimation with missing data in latent factor models.

By Raaz Dwivedi, Sabina Tomkins, Predrag Klasnja, Susan Murphy, Devavrat Shah
arXiv Machine Learning
Sep 1

AI-Generated Measurements for Identification and Inference with Missing Data: A Weak Shadow Variable Approach

The paper introduces an assumption‑lean framework that uses AI‑generated measurements as weak shadow variables to identify and infer population quantities when data are missing not at random. Weak shadow variables are outcome‑informative proxies that are conditionally independent of missingness given the true outcome and covariates, and they do not need to predict missing outcomes accurately. The authors derive sharp bounds via linear programs and propose a localized penalized estimator with a subsampling algorithm for confidence intervals, demonstrating in semi‑synthetic experiments that the resulting intervals are substantially narrower and more accurate than classical MNAR methods.

By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv Machine Learning
Jun 9

Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models

arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.

By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv Machine Learning
Jul 15

Linear Regression under Missing or Corrupted Coordinates

arXiv:2509. 19242v2 Announce Type: replace-cross Abstract: We study multivariate linear regression under Gaussian covariates in two settings, where data may be erased or corrupted by an adversary under a coordinate-wise budget.

By Ilias Diakonikolas, Jelena Diakonikolas, Daniel M. Kane, Jasper C. H. Lee, Thanasis Pittas
arXiv Machine Learning
Sep 10

Low-Rank Plus Sparse Matrix Transfer Learning under Growing Representations and Ambient Dimensions

The paper introduces a transfer learning framework for structured matrix estimation when both the ambient dimension and the intrinsic representation grow over time. It models the target parameter as an embedded source component plus low‑rank innovations and sparse edits, and proposes an anchored alternating projection estimator that preserves the transferred subspace while estimating only the new components. Deterministic error bounds are derived that separate target noise, representation growth, and source estimation error, showing improved rates when rank and sparsity increments are small, and the framework is applied to Markov transition matrix estimation and structured covariance estimation with theoretical guarantees and empirical validation.

By Jinhang Chai, Xuyuan Liu, Elynn Chen, Yujun Yan