arXiv Machine Learning

Learning Who to Treat When Treatment is Missing

arXiv:2607. 14346v1 Announce Type: new Abstract: Policy learning methods are increasingly used to inform treatment allocation under budget constraints.

arXiv Statistics ML
Sep 22

Doubly robust target inference for generalized linear regression with completely missing covariates

arXiv:2609.24086v1 Announce Type: cross Abstract: Large-scale multipurpose cohort studies and biobanks often omit covariates needed for specific downstream analyses. We study target-population infere...

By Huali Zhao (School of Mathematics and Statistics, Huazhong University of Science and Technology), Ke Deng (Department of Statistics and Data Science, Tsinghua University)
arXiv Machine Learning
Jun 9

Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models

arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.

By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv Machine Learning
Sep 18

Stable Policy Learning

The paper investigates how policy learning algorithms should balance expected welfare against sampling risk in evidence-based policymaking. It demonstrates that algorithmic stability—specifically, a policy’s insensitivity to the replacement of a single experimental unit—limits sampling risk. The authors introduce policy‑vote bagging, which trains on many subsamples and averages their votes, preserving expected welfare while improving expected utility for risk‑averse researchers, and provide sharp bounds linking estimation accuracy, subsample size, and welfare variation, including an exact guarantee under CARA utility.

By Harvey Barnhard, Giacomo Opocher, Rahul Singh
arXiv Machine Learning
Sep 14

MInTRL: Off-policy Intervention can boost On-policy RL

MInTRL (Minimal Intervention Reinforcement Learning) expands exploration in on-policy reinforcement learning by inserting sparse, local corrections into rollouts via a judge-intervention policy. These interventions replace erroneous suffixes and immediately return control to the main policy, allowing the agent to explore beyond its natural trajectory while maintaining on-policy data. The method uses a sequence-level advantage-regression objective, avoiding importance sampling, and demonstrates superior performance on math and code benchmarks compared to standard on-policy and off-policy baselines.

By Mingyu Chen, Yefan Tao, Gerald Friedland, Xuezhou Zhang, Chris Kong