arXiv:2602. 17894v2 Announce Type: replace-cross Abstract: Data collection is a critical component of modern statistical and machine learning pipelines, particularly when data must be gathered from multiple heterogeneous sources to study a target population of interest.
By Michael O. Harding, Vikas Singh, Kirthevasan Kandasamy
The paper introduces a model‑agnostic inference framework for partially identified causal effects that leverages covariate information without requiring discrete covariates or accurate conditional distribution estimates. Using duality theory for optimal transport, the method delivers uniformly valid inference in randomized experiments, is doubly robust in observational settings, achieves asymptotic unbiasedness when nuisance parameters converge semiparametrically, and allows multiplier‑bootstrap selection of covariates and models while remaining computationally efficient. Empirical applications show the approach consistently narrows identified sets and confidence intervals without imposing extra structural assumptions.
By Wenlong Ji, Lihua Lei, Asher Spector
The paper investigates preference elicitation under the Bradley‑Terry‑Luce model, focusing on estimating an unknown partworth vector from pairwise queries that satisfy a joint identifiability condition. It derives minimax lower bounds and shows that the canonical maximum likelihood estimator (MLE) exists, is unique, and achieves near‑optimal error rates once the sample size exceeds a design‑dependent threshold, without requiring compactness constraints or external regularizers. The analysis decomposes the estimation error into a linear stochastic term, a second‑order bias, and a higher‑order remainder, providing a unified non‑asymptotic theory for parametric utility elicitation.
By Yicheng Li, Huifu Xu
The paper investigates Double Machine Learning (DML) estimators under structure‑agnostic (SA) models, which assume the data‑generating law lies within a neighborhood of fixed machine‑learning estimates. It shows that for two of three studied functionals—the quadratic functional in the Gaussian sequence model and the quadratic density integral functional—the DML estimators are asymptotically inadmissible, being dominated by second‑order empirical higher‑order influence function (HOIF) estimators. For the third functional, the expected conditional covariance, both DML and HOIF estimators remain minimax but neither dominates the other.
By Lin Liu, Rajarshi Mukherjee, James M Robins
arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.
By Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
arXiv:2603. 17925v2 Announce Type: replace-cross Abstract: We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms.
By Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan
arXiv:2607. 06879v1 Announce Type: new Abstract: Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly.
By Tianyi Ma, Hanzhang Qin, Ruihao Zhu, Jierui Zuo
arXiv:2604. 14575v3 Announce Type: replace-cross Abstract: Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decisions, and field experiment outcomes.
By Cheng Lu, Mengxin Wang, Dennis J. Zhang, Heng Zhang
arXiv:2606. 15600v1 Announce Type: cross Abstract: Cardinality-estimation (CE) research ranks estimators by q-error, yet it is well known that q-error is an imperfect proxy for query-plan quality.
By Madhulatha Mandarapu, Sandeep Kunkunuru
arXiv:2606. 00563v1 Announce Type: cross Abstract: Selection bias is a common and often unavoidable aspect of real-world data that challenges the generalizability of machine learning models.
By Kara Liu, Maggie Wang, Russ B. Altman
arXiv:2602. 04402v3 Announce Type: replace-cross Abstract: Performative predictions influence the very outcomes they aim to forecast.
By Julian Rodemann, Unai Fischer-Abaigar, James Bailie, Krikamol Muandet
arXiv:2609.06294v1 Announce Type: new
Abstract: Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental s...
By Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez, Tamar Krishnamurti, Bryan Wilder