The paper introduces a new algorithm that uses decision trees and random forests to estimate individual treatment effects while providing interpretability. It modifies the standard random forest splitting criterion by combining a heterogeneity-focused criterion with a bias-correction criterion, enabling the model to handle observational studies with varying treatment propensities without separately estimating propensity scores. The resulting tree structure directly reveals which features drive treatment effect differences, and simulation studies show the method matches or surpasses existing approaches in prediction accuracy while improving interpretability.
By Nicolas Alexander Ihlo, Merle Behr
arXiv:2604. 23904v3 Announce Type: replace-cross Abstract: Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference.
By Yichen Xu
arXiv:2609.35815v1 Announce Type: cross
Abstract: Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated co...
By Ian Arawjo
arXiv:2606. 23880v1 Announce Type: new Abstract: From climate teleconnections to gene regulation, modern time-series datasets encompass tens or hundreds of interacting variables, making causal discovery increasingly challenging.
By Mohammad Fesanghary, Abhinav Havaldar
The paper introduces a method for selecting the best heterogeneous treatment effect (HTE) estimator from a set of candidates when the true treatment effect is unobserved. It frames estimator selection as a multiple testing problem and proposes a cross‑fitted, exponentially weighted test statistic that uses a two‑way sample splitting scheme to separate nuisance estimation from weight learning, ensuring stability for inference. The authors prove asymptotic familywise error rate control under mild conditions and demonstrate empirically that their procedure reduces false selections compared to common methods on ACIC 2016, IHDP, and Twins benchmarks.
By Jiayi Guo, Zijun Gao
The paper introduces GeoACE, a five‑expert framework for estimating heterogeneous treatment effects that blends a common anchor‑correction estimator with overlap‑aware and outcome‑guided geometries. The ensemble’s task‑level weights are learned from internal validation predictions, frozen before test evaluation, and applied to experts refitted on the full development data. Adding the outcome‑free, overlap‑aware expert O‑Phi‑ACE consistently improves performance across seven benchmarks, achieving the lowest average rank among 11 comparators.
By Ali Haghpanah Jahromi, Mohammad Taheri
arXiv:2603. 19186v3 Announce Type: replace Abstract: Randomized controlled trials (RCTs) are the gold standard for estimating treatment effects, yet they are often underpowered for detecting effect heterogeneity.
By Amir Asiaee, Samhita Pal
arXiv:2607. 23721v1 Announce Type: cross Abstract: Distributional random forests replace mean-based CART splitting with criteria that compare the full conditional response distribution in candidate children.
By Silas Koemen
arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.
By Yilin Song, Dan M. Kluger, Harsh Parikh, Tian Gu
arXiv:2606. 20820v2 Announce Type: replace Abstract: Can we trust evaluation scores to capture an LLM's true real-world performance?
By Zhijian Zhou, Zesheng Ye, Zhaorun Chen, Bo Li, Feng Liu
arXiv:2609.26142v1 Announce Type: cross
Abstract: Augmented inverse-probability weighting (AIPW), targeted maximum likelihood estimation (TMLE), and double/debiased machine learning (DML) are three r...
By M. Ehsan Karim
arXiv:2609.06294v1 Announce Type: new
Abstract: Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental s...
By Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez, Tamar Krishnamurti, Bryan Wilder