arXiv Machine Learning

Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence

arXiv:2607. 29456v1 Announce Type: cross Abstract: Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine learning methods to be used for nuisance parameter estimation while preserving valid inference.

arXiv Machine Learning
Aug 11

Demystifying Prediction Powered Inference

arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.

By Yilin Song, Dan M. Kluger, Harsh Parikh, Tian Gu
arXiv Machine Learning
Sep 10

Conformalized Super Learner

arXiv:2604.22391v2 Announce Type: replace-cross Abstract: The Super Learner (SL) is a widely used ensemble method that combines point predictions from a library of learners based on their predictive...

By Zhanli Wu, Fabrizio Leisen, Miguel-Angel Luque-Fernandez, F. Javier Rubio
arXiv Statistics ML
Sep 7

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models

The paper investigates Double Machine Learning (DML) estimators under structure‑agnostic (SA) models, which assume the data‑generating law lies within a neighborhood of fixed machine‑learning estimates. It shows that for two of three studied functionals—the quadratic functional in the Gaussian sequence model and the quadratic density integral functional—the DML estimators are asymptotically inadmissible, being dominated by second‑order empirical higher‑order influence function (HOIF) estimators. For the third functional, the expected conditional covariance, both DML and HOIF estimators remain minimax but neither dominates the other.

By Lin Liu, Rajarshi Mukherjee, James M Robins
arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto
arXiv Machine Learning
Sep 16

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias

The paper introduces a new algorithm that uses decision trees and random forests to estimate individual treatment effects while providing interpretability. It modifies the standard random forest splitting criterion by combining a heterogeneity-focused criterion with a bias-correction criterion, enabling the model to handle observational studies with varying treatment propensities without separately estimating propensity scores. The resulting tree structure directly reveals which features drive treatment effect differences, and simulation studies show the method matches or surpasses existing approaches in prediction accuracy while improving interpretability.

By Nicolas Alexander Ihlo, Merle Behr