arXiv Machine Learning

AI-Generated Measurements for Identification and Inference with Missing Data: A Weak Shadow Variable Approach

The paper introduces an assumption‑lean framework that uses AI‑generated measurements as weak shadow variables to identify and infer population quantities when data are missing not at random. Weak shadow variables are outcome‑informative proxies that are conditionally independent of missingness given the true outcome and covariates, and they do not need to predict missing outcomes accurately. The authors derive sharp bounds via linear programs and propose a localized penalized estimator with a subsampling algorithm for confidence intervals, demonstrating in semi‑synthetic experiments that the resulting intervals are substantially narrower and more accurate than classical MNAR methods.

arXiv Machine Learning
Jun 9

Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models

arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.

By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv AI
Aug 20

Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements

The paper introduces Debiased Inference with Multiple Imperfect Measurements (DMM), a framework that uses several error‑prone AI measurements to perform valid downstream statistical inference without requiring costly gold‑standard labels. By assuming conditional independence of the measurements given the true label and unit‑level features, DMM leverages CP decomposition and semiparametric theory to prove consistency and asymptotic normality of its estimator. Simulations demonstrate that DMM yields valid inference and can improve efficiency when additional imperfect measurements are available, and the authors provide diagnostics for the key independence assumption.

By Naoki Egami, Sooahn Shin
arXiv Machine Learning
4d ago

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective

The paper discusses how large language models (LLMs) can be fine‑tuned with observational data to improve alignment with human preferences and business goals. It highlights that directly using such data can cause models to learn spurious correlations, and introduces DeconfoundLM, a method that removes known confounders from reward signals. Experiments show that DeconfoundLM better recovers causal relationships and outperforms baseline methods by over 16% in objective score when confounding is present.

By Erfan Loghmani
arXiv AI
Jun 4

Generative Augmented Inference

arXiv:2604. 14575v2 Announce Type: replace-cross Abstract: Large language models enable inexpensive AI-generated annotations, but using them reliably for causal inference remains challenging.

By Cheng Lu, Mengxin Wang, Dennis J. Zhang, Heng Zhang
arXiv Machine Learning
Jul 2

Deep learning with missing data

arXiv:2504. 15388v3 Announce Type: replace-cross Abstract: In the context of multivariate nonparametric regression with missing covariates, we propose Pattern Embedded Neural Networks (PENNs), which can be applied in conjunction with any existing imputation technique.

By Tianyi Ma, Tengyao Wang, Richard J. Samworth