Large language models have been widely evaluated as simulators of individual survey responses. In practice, however, fully unobserved responses are rare; the dominant problem is partial non-response.
arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.
By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
The paper introduces an assumption‑lean framework that uses AI‑generated measurements as weak shadow variables to identify and infer population quantities when data are missing not at random. Weak shadow variables are outcome‑informative proxies that are conditionally independent of missingness given the true outcome and covariates, and they do not need to predict missing outcomes accurately. The authors derive sharp bounds via linear programs and propose a localized penalized estimator with a subsampling algorithm for confidence intervals, demonstrating in semi‑synthetic experiments that the resulting intervals are substantially narrower and more accurate than classical MNAR methods.
By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv:2607. 07767v1 Announce Type: cross Abstract: Missing values undermine statistical inference and machine learning pipelines, yet most imputation methods rely on heuristics or restrictive parametric assumptions that ignore the joint data distribution.
By Andrea Basteri, Carlo Ciliberto, Alessandro Rudi
The paper presents an uncertainty‑aware machine‑learning approach for mapping poverty in Africa using satellite imagery. By combining simultaneous quantile regression with a novel conformal prediction technique, the authors generate statistically guaranteed prediction intervals for neighborhood‑level International Wealth Index estimates, achieving high explanatory power (R² = 0.75) while acknowledging broader uncertainty. They also propose a risk‑controlled aid allocation procedure that leverages both survey data and model predictions, showing in simulations that it can deliver more aid per eligible recipient than alternative strategies.
By Markus B. Pettersson, James Bailie, Mohammad Kakooei, Eagon Meng, Adel Daoud
arXiv:2507. 03897v3 Announce Type: replace Abstract: We introduce GenAI-Powered Inference (GPI), a statistical framework for both causal and predictive inference using unstructured data, including text and images.
By Kosuke Imai, Kentaro Nakamura
The paper introduces PSL, a dual‑view framework that enhances large language models for predictive political question answering by leveraging semi‑structured political records. PSL extracts stance signals from actor records in a semantic view and learns structure‑aware actor representations from an interaction graph in a vector view. Experiments on three real‑world datasets show that PSL consistently outperforms baseline methods, with ablation studies confirming the complementary benefits of stance and structure signals.
By Yinan Liu, Zihan Zhou, Zichun Jin, Xinyu Wang, Bin Wang, Xiaochun Yang
arXiv:2606. 20538v1 Announce Type: new Abstract: Bayesian predictive inference provides a principled framework for uncertainty quantification, data efficiency, and robust generalization.
By Qingyang Zhu, Eric Karl Oermann, Kyunghyun Cho
arXiv:2606. 03347v1 Announce Type: cross Abstract: Score-based diffusion models have emerged as prominent deep generative models; however, their application to tabular data remains challenging because their backbones assume fully specified inputs, whereas real-world tabular data often contain missing values.
By Jungkyu Kim, Taeyoung Park, Kibok Lee
arXiv:2604. 07635v2 Announce Type: replace-cross Abstract: This research considers a scalable inference for spatial data modeled through Gaussian intrinsic conditional autoregressive (ICAR) structures.
By Debjoy Thakur
arXiv:2608. 15121v1 Announce Type: cross Abstract: Sufficient dimension reduction (SDR) seeks the minimal subspace of the predictors that captures the full conditional distribution of the response, which is known as the central subspace (CS).
By Ye Tian
arXiv:2507. 03897v4 Announce Type: replace Abstract: We introduce GenAI-Powered Inference (GPI), a statistical framework for both causal and predictive inference using unstructured data, including text and images.
By Kosuke Imai, Kentaro Nakamura