The paper introduces Debiased Inference with Multiple Imperfect Measurements (DMM), a framework that uses several error‑prone AI measurements to perform valid downstream statistical inference without requiring costly gold‑standard labels. By assuming conditional independence of the measurements given the true label and unit‑level features, DMM leverages CP decomposition and semiparametric theory to prove consistency and asymptotic normality of its estimator. Simulations demonstrate that DMM yields valid inference and can improve efficiency when additional imperfect measurements are available, and the authors provide diagnostics for the key independence assumption.
By Naoki Egami, Sooahn Shin
The paper introduces AICOME, a framework that uses AI-derived respondent-level measures to recover both individual and group-level effects in contextual models. By applying AICOME to the 2022 China Family Panel Studies, the authors validate that AI measures can replicate key survey variables such as computer use, foreign-language use, weekly hours, and management responsibilities, especially when rich respondent and job data are available. The study also identifies boundary conditions, noting reduced performance when only occupation and basic demographics are used or when multiple related constructs are simultaneously unobserved.
By Wenxin Jiang, Xuyang Wang, Yuxiao Wu
arXiv:2606. 21185v2 Announce Type: replace-cross Abstract: There is a precise sense in which drawing causal inferences from observational data is hard, even when identifiability is assumed.
By Alexis Bellot
The paper introduces an assumption‑lean framework that uses AI‑generated measurements as weak shadow variables to identify and infer population quantities when data are missing not at random. Weak shadow variables are outcome‑informative proxies that are conditionally independent of missingness given the true outcome and covariates, and they do not need to predict missing outcomes accurately. The authors derive sharp bounds via linear programs and propose a localized penalized estimator with a subsampling algorithm for confidence intervals, demonstrating in semi‑synthetic experiments that the resulting intervals are substantially narrower and more accurate than classical MNAR methods.
By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv:2411.09686v4 Announce Type: replace
Abstract: Regressing a function $F$ on $\mathbb{R}^d$ without incurring the statistical and computational curse of dimensionality requires exploitable struct...
By Yantao Wu, Mauro Maggioni
arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.
By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong