arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.
By Yilin Song, Dan M. Kluger, Harsh Parikh, Tian Gu
arXiv:2607. 28655v1 Announce Type: cross Abstract: Small-area estimation (SAE) enables researchers and policymakers to identify spatial disparities in health outcomes, but survey-based SAE products carry an inherent lag.
By Aanya Gupta, Szandra P\'eter, Sara Von Hoene, Emma Von Hoene, Taylor Anderson
arXiv:2609.26142v1 Announce Type: cross
Abstract: Augmented inverse-probability weighting (AIPW), targeted maximum likelihood estimation (TMLE), and double/debiased machine learning (DML) are three r...
By M. Ehsan Karim
arXiv:2604.22391v2 Announce Type: replace-cross
Abstract: The Super Learner (SL) is a widely used ensemble method that combines point predictions from a library of learners based on their predictive...
By Zhanli Wu, Fabrizio Leisen, Miguel-Angel Luque-Fernandez, F. Javier Rubio
arXiv:2609.17238v1 Announce Type: cross
Abstract: High-dimensional data create challenges for causal effect estimation because identifying the covariates needed for correct model specification become...
By Muwon Kwon, Peter M. Steiner
The paper investigates Double Machine Learning (DML) estimators under structure‑agnostic (SA) models, which assume the data‑generating law lies within a neighborhood of fixed machine‑learning estimates. It shows that for two of three studied functionals—the quadratic functional in the Gaussian sequence model and the quadratic density integral functional—the DML estimators are asymptotically inadmissible, being dominated by second‑order empirical higher‑order influence function (HOIF) estimators. For the third functional, the expected conditional covariance, both DML and HOIF estimators remain minimax but neither dominates the other.
By Lin Liu, Rajarshi Mukherjee, James M Robins
arXiv:2606. 00563v1 Announce Type: cross Abstract: Selection bias is a common and often unavoidable aspect of real-world data that challenges the generalizability of machine learning models.
By Kara Liu, Maggie Wang, Russ B. Altman
arXiv:2411.02771v3 Announce Type: replace-cross
Abstract: Doubly robust estimators are widely used for estimating average treatment effects and other linear summaries of regression functions. While c...
By Lars van der Laan, Alex Luedtke, Marco Carone
arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.
By Diego Marcondes, Cl\'audia Peixoto
The paper introduces a new algorithm that uses decision trees and random forests to estimate individual treatment effects while providing interpretability. It modifies the standard random forest splitting criterion by combining a heterogeneity-focused criterion with a bias-correction criterion, enabling the model to handle observational studies with varying treatment propensities without separately estimating propensity scores. The resulting tree structure directly reveals which features drive treatment effect differences, and simulation studies show the method matches or surpasses existing approaches in prediction accuracy while improving interpretability.
By Nicolas Alexander Ihlo, Merle Behr
arXiv:2109.02355v2 Announce Type: replace
Abstract: The last decade of progress in machine learning (ML), especially the deep learning era, has raised a number of scientific questions that challenge...
By Yehuda Dar, Vidya Muthukumar, Richard G. Baraniuk
arXiv:2606. 09860v1 Announce Type: cross Abstract: Non-alcoholic fatty liver disease (NAFLD) affects roughly 25% of global adults, posing substantial hepatic and cardiovascular risks.
By Xinze Zhang