arXiv Machine Learning

Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models

arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.

arXiv Machine Learning
Sep 1

AI-Generated Measurements for Identification and Inference with Missing Data: A Weak Shadow Variable Approach

The paper introduces an assumption‑lean framework that uses AI‑generated measurements as weak shadow variables to identify and infer population quantities when data are missing not at random. Weak shadow variables are outcome‑informative proxies that are conditionally independent of missingness given the true outcome and covariates, and they do not need to predict missing outcomes accurately. The authors derive sharp bounds via linear programs and propose a localized penalized estimator with a subsampling algorithm for confidence intervals, demonstrating in semi‑synthetic experiments that the resulting intervals are substantially narrower and more accurate than classical MNAR methods.

By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv AI
Aug 20

Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements

The paper introduces Debiased Inference with Multiple Imperfect Measurements (DMM), a framework that uses several error‑prone AI measurements to perform valid downstream statistical inference without requiring costly gold‑standard labels. By assuming conditional independence of the measurements given the true label and unit‑level features, DMM leverages CP decomposition and semiparametric theory to prove consistency and asymptotic normality of its estimator. Simulations demonstrate that DMM yields valid inference and can improve efficiency when additional imperfect measurements are available, and the authors provide diagnostics for the key independence assumption.

By Naoki Egami, Sooahn Shin
arXiv Machine Learning
Jul 7

How Much is Left? LLMs Linearly Encode Their Remaining Output Length

arXiv:2607. 05316v1 Announce Type: cross Abstract: Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts.

By Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi, Damiano Fornasiere, Adam Oberman
arXiv AI
Aug 12

On Solomonoff Induction in Large Language Models and the Limits of Self-Improving: The Singularity Is Not Near Without Symbolic Model Synthesis

arXiv:2601. 05280v3 Announce Type: replace-cross Abstract: On the one hand, the question of whether large language models (LLMs) are Solomonoff induction estimators has become an explicit question at the intersection of Algorithmic Information Theory (AIT) and Machine Learning (ML) of great interest.

By Hector Zenil
arXiv AI
Sep 21

Large Language Models As Shannon Lossy Compressors Not Solomonoff Induction Estimators: The Singularity Is Not Near Without Symbolic Model Synthesis

The paper argues that Large Language Models (LLMs) do not function as Solomonoff induction estimators because their training objectives—cross‑entropy, negative log‑likelihood, and next‑token prediction—optimize fit to a supplied conditional distribution rather than a program‑weighted universal mixture. It further contends that additional computation alone does not transform these models into optimal predictors without external hyper‑parameter or architectural changes. The authors suggest that neurosymbolic machine learning, exemplified by models such as Fable and Astra, represents a shift toward symbolic model synthesis, moving beyond purely statistical LLMs.

By Hector Zenil, Abicumaran Uthamacumaran, Luan Ozelim