arXiv Statistics ML

Three Routes to One Answer: Reconciling AIPW, TMLE, and Double Machine Learning for Applied Researchers

arXiv Machine Learning
Jun 4

Orthogonal Learner for Estimating Heterogeneous Long-Term Treatment Effects

arXiv:2604. 00915v2 Announce Type: replace Abstract: Estimation of heterogeneous long-term treatment effects (HLTEs) is relevant for personalized decision-making in marketing, economics, and medicine, where short-term observational datasets are often combined with long-term observational datasets.

By Haorui Ma, Dennis Frauen, Valentyn Melnychuk, Stefan Feuerriegel
arXiv Machine Learning
Aug 3

Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence

arXiv:2607. 29456v1 Announce Type: cross Abstract: Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine learning methods to be used for nuisance parameter estimation while preserving valid inference.

By Haozheng Xu, Siyuan Ma, Qingyan Xiang
arXiv AI
Jul 8

K-ABENA: K-Adaptive Backpropagation with Error-based N-exclusion Algorithm : (Compensated Loss-Based Sample Exclusion with Unbiased Gradient Estimation)

arXiv:2607. 05903v1 Announce Type: cross Abstract: We present K-ABENA (K-Adaptive Backpropagation with Error-based N-exclusion Algorithm), a selective gradient computation framework that reduces per-iteration training cost by excluding a fraction of low-loss ("minor") observations from the backward pass.

By Jean-Francois Bonbhel
arXiv Statistics ML
Aug 24

Double Machine Learning of Continuous Treatment Effects with Additive Instrumental Variables

The paper introduces a new framework for identifying average dose-response functions in the presence of unmeasured confounding by using instrumental variables. It defines a uniform regular weighting function and partitions the treatment space into open sets where local identification is possible. For estimation, the authors propose an augmented inverse probability weighted score within a debiased machine learning setting, along with practical guidance for constructing weighting functions, falsification tests for the additive IV condition, and asymptotic theory for kernel regression or empirical risk minimization estimators.

By Shuyuan Chen, Peng Zhang, Yifan Cui
arXiv Machine Learning
Sep 4

Reliable Selection of Heterogeneous Treatment Effect Estimators

The paper introduces a method for selecting the best heterogeneous treatment effect (HTE) estimator from a set of candidates when the true treatment effect is unobserved. It frames estimator selection as a multiple testing problem and proposes a cross‑fitted, exponentially weighted test statistic that uses a two‑way sample splitting scheme to separate nuisance estimation from weight learning, ensuring stability for inference. The authors prove asymptotic familywise error rate control under mild conditions and demonstrate empirically that their procedure reduces false selections compared to common methods on ACIC 2016, IHDP, and Twins benchmarks.

By Jiayi Guo, Zijun Gao
arXiv Statistics ML
Sep 7

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models

The paper investigates Double Machine Learning (DML) estimators under structure‑agnostic (SA) models, which assume the data‑generating law lies within a neighborhood of fixed machine‑learning estimates. It shows that for two of three studied functionals—the quadratic functional in the Gaussian sequence model and the quadratic density integral functional—the DML estimators are asymptotically inadmissible, being dominated by second‑order empirical higher‑order influence function (HOIF) estimators. For the third functional, the expected conditional covariance, both DML and HOIF estimators remain minimax but neither dominates the other.

By Lin Liu, Rajarshi Mukherjee, James M Robins
arXiv Machine Learning
Sep 16

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias

The paper introduces a new algorithm that uses decision trees and random forests to estimate individual treatment effects while providing interpretability. It modifies the standard random forest splitting criterion by combining a heterogeneity-focused criterion with a bias-correction criterion, enabling the model to handle observational studies with varying treatment propensities without separately estimating propensity scores. The resulting tree structure directly reveals which features drive treatment effect differences, and simulation studies show the method matches or surpasses existing approaches in prediction accuracy while improving interpretability.

By Nicolas Alexander Ihlo, Merle Behr