arXiv Statistics ML

Learning-Enabled Estimation: Tight Characterizations under Sample Selection Biases

The paper investigates regression when outcomes are observed only after passing through selection filters that depend on both covariates and outcomes, a common issue in fields such as clinical trials, labor markets, and auctions. It provides a complete characterization of the minimal assumptions on the functional forms of selection processes that allow regression to remain possible, and shows that the regression function can sometimes be identified even when the selection filter itself cannot. Under stronger identification conditions, the authors also deliver finite‑sample estimation guarantees, explicit convergence rates, and oracle‑efficient algorithms, offering the first general‑purpose estimation method for this broad class of selection problems.

arXiv Statistics ML
Aug 25

Model-Agnostic Covariate-Assisted Inference on Partially Identified Causal Effects

The paper introduces a model‑agnostic inference framework for partially identified causal effects that leverages covariate information without requiring discrete covariates or accurate conditional distribution estimates. Using duality theory for optimal transport, the method delivers uniformly valid inference in randomized experiments, is doubly robust in observational settings, achieves asymptotic unbiasedness when nuisance parameters converge semiparametrically, and allows multiplier‑bootstrap selection of covariates and models while remaining computationally efficient. Empirical applications show the approach consistently narrows identified sets and confidence intervals without imposing extra structural assumptions.

By Wenlong Ji, Lihua Lei, Asher Spector
arXiv Machine Learning
Sep 23

Error Bounds for Statistical Estimators in BTL Model with Parametric Multivariate Utility Functions

The paper investigates preference elicitation under the Bradley‑Terry‑Luce model, focusing on estimating an unknown partworth vector from pairwise queries that satisfy a joint identifiability condition. It derives minimax lower bounds and shows that the canonical maximum likelihood estimator (MLE) exists, is unique, and achieves near‑optimal error rates once the sample size exceeds a design‑dependent threshold, without requiring compactness constraints or external regularizers. The analysis decomposes the estimation error into a linear stochastic term, a second‑order bias, and a higher‑order remainder, providing a unified non‑asymptotic theory for parametric utility elicitation.

By Yicheng Li, Huifu Xu
arXiv Statistics ML
Sep 7

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models

The paper investigates Double Machine Learning (DML) estimators under structure‑agnostic (SA) models, which assume the data‑generating law lies within a neighborhood of fixed machine‑learning estimates. It shows that for two of three studied functionals—the quadratic functional in the Gaussian sequence model and the quadratic density integral functional—the DML estimators are asymptotically inadmissible, being dominated by second‑order empirical higher‑order influence function (HOIF) estimators. For the third functional, the expected conditional covariance, both DML and HOIF estimators remain minimax but neither dominates the other.

By Lin Liu, Rajarshi Mukherjee, James M Robins
arXiv Machine Learning
Jun 9

Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models

arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.

By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv Machine Learning
Jun 5

Multi-Armed Sequential Hypothesis Testing by Betting

arXiv:2603. 17925v2 Announce Type: replace-cross Abstract: We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms.

By Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan
arXiv Machine Learning
Jul 9

Best-Arm Identification with Generative Proxy

arXiv:2607. 06879v1 Announce Type: new Abstract: Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly.

By Tianyi Ma, Hanzhang Qin, Ruihao Zhu, Jierui Zuo
arXiv AI
Jun 9

Performative Learning Theory

arXiv:2602. 04402v3 Announce Type: replace-cross Abstract: Performative predictions influence the very outcomes they aim to forecast.

By Julian Rodemann, Unai Fischer-Abaigar, James Bailie, Krikamol Muandet