Hugging Face Trending Papers

In-Context Learning for the Imputation of Public Opinion Data with Large Language Models

Read the original on Hugging Face Trending Papers →

Large language models have been widely evaluated as simulators of individual survey responses. In practice, however, fully unobserved responses are rare; the dominant problem is partial non-response.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 1

AI-Generated Measurements for Identification and Inference with Missing Data: A Weak Shadow Variable Approach

The paper introduces an assumption‑lean framework that uses AI‑generated measurements as weak shadow variables to identify and infer population quantities when data are missing not at random. Weak shadow variables are outcome‑informative proxies that are conditionally independent of missingness given the true outcome and covariates, and they do not need to predict missing outcomes accurately. The authors derive sharp bounds via linear programs and propose a localized penalized estimator with a subsampling algorithm for confidence intervals, demonstrating in semi‑synthetic experiments that the resulting intervals are substantially narrower and more accurate than classical MNAR methods.

By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv Machine Learning
Jun 9

Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models

arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.

By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong