Semiparametric Inference for Conditional Shapley Feature Importance
Read the original on arXiv Statistics ML →The Flow has not summarised this story yet — read it at arXiv Statistics ML.
The Flow has not summarised this story yet — read it at arXiv Statistics ML.
arXiv:2604. 15107v2 Announce Type: replace-cross Abstract: Shapley values provide a flexible framework for attributing feature contributions to model predictions, but they are not naturally suited for feature selection: a feature may receive a positive attribution even when it is redundant given the remaining variables.
arXiv:2607. 05806v1 Announce Type: new Abstract: Training data for machine learning is routinely collected by a selection process the model never sees: loans are observed only when granted, outcomes only when a test was ordered.
arXiv:2607. 23721v1 Announce Type: cross Abstract: Distributional random forests replace mean-based CART splitting with criteria that compare the full conditional response distribution in candidate children.
arXiv:2609.14902v1 Announce Type: cross Abstract: Shapley value (SV)-based methods are the prevailing framework for feature attribution in machine learning, yet existing population-level Shapley esti...
arXiv:2610.01641v1 Announce Type: cross Abstract: Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive...
The paper introduces a model‑agnostic inference framework for partially identified causal effects that leverages covariate information without requiring discrete covariates or accurate conditional distribution estimates. Using duality theory for optimal transport, the method delivers uniformly valid inference in randomized experiments, is doubly robust in observational settings, achieves asymptotic unbiasedness when nuisance parameters converge semiparametrically, and allows multiplier‑bootstrap selection of covariates and models while remaining computationally efficient. Empirical applications show the approach consistently narrows identified sets and confidence intervals without imposing extra structural assumptions.