arXiv Machine Learning

Exact Functional ANOVA Decomposition for Categorical Inputs Models

arXiv:2603. 02673v2 Announce Type: replace-cross Abstract: Functional ANOVA offers a principled framework for interpretability by decomposing a model's prediction into main effects and higher-order interactions.

arXiv Machine Learning
Sep 18

Evaluating Explanation Methods by the Predictors They Induce

The paper proposes a new evaluation test for explanation methods: if an explanation accurately captures how a model uses its features, one should be able to reconstruct the model’s predictions from it. The authors convert explanations into predictors by summing feature effects and assess how well these predictors reproduce the model on unseen data, without any fitting. They apply this test to partial dependence plots, accumulated local effects, SHAP, and LIME across multiple datasets and model families, showing that the best method depends on feature dependence and that some existing quality metrics can favor flawed explanations.

By Jacob Selb{\ae}k, Hugo L. Hammer
arXiv Machine Learning
Sep 17

NObSP: Functional Decomposition of Neural Networks via Oblique Subspace Projections

NObSP (Nonlinear Oblique Subspace Projections) is a framework that decomposes neural network predictions into explicit per‑feature contribution functions and an interaction residual, leveraging the linear final layer and oblique projections to avoid double counting when feature subspaces overlap. It connects to functional ANOVA and the Kolmogorov‑Arnold representation theorem, and introduces an efficient partial regression algorithm for out‑of‑sample evaluation. For convolutional networks, NObSP‑CAM generates class activation maps without backward passes after a single calibration, and experiments on tabular and vision datasets show faithfulness comparable to established attribution methods, with high function reproduction scores and improved class purity on TinyImageNet.

By Alexander Caicedo, V\'ictor De La Hoz, Santiago Alf\'erez
arXiv AI
Sep 16

A unified framework for global and local interpretability using adaptive derivative-ordered random explanation

The paper introduces Adaptive Derivative-Ordered Random Explanation (ADORE), a unified framework that uses first- and second-order derivatives to capture nonlinear feature interactions and feature-sample dynamics. ADORE combines global feature importance with local sample contributions, quantifying both magnitude and direction of feature impact while identifying critical samples. It achieves computational efficiency via randomized SVD and dynamic sparsity detection, outperforming LIME and SHAP across tabular, text, and image data, and is released as an open-source Python package on GitHub.

By Lemen Chao, Ming Lei, Anran Fanga
arXiv Machine Learning
Jun 9

Disentangled Feature Importance

arXiv:2507. 00260v3 Announce Type: replace-cross Abstract: When predictors are statistically dependent, the appropriate definition of feature importance depends on the operational goal.

By Jin-Hong Du, Kathryn Roeder, Larry Wasserman
arXiv AI
Aug 24

GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series

GroupSegment-SHAP (GS‑SHAP) introduces explanatory units called group‑segment players that capture cross‑variable dependence and distribution shifts over time in multivariate time‑series models. By attributing Shapley values to these units, GS‑SHAP preserves joint structural signals that traditional time‑series SHAP variants fragment. Experiments on human activity recognition, power‑system forecasting, medical signal analysis, and financial time series show that GS‑SHAP improves deletion‑based faithfulness by about 1.7× and reduces runtime by roughly 40% compared to existing baselines, while a financial case study demonstrates its ability to reveal interpretable multivariate‑temporal interactions during high‑volatility periods.

By Jinwoong Kim, Sangjin Park
arXiv Machine Learning
Aug 24

Conditional-Independence-Regularized Distributional Autoencoders for Mixed-Type Data

The paper introduces Conditional-Independence-Regularized Distributional Autoencoders, a framework for learning low-dimensional representations of mixed-type data that includes both numerical and categorical variables. It uses an energy-score objective for numerical variables, a likelihood objective for categorical variables, and an auxiliary conditional independence regularization term to capture dependencies between variable types. The authors provide theoretical analysis and demonstrate that the method improves categorical distribution recovery, achieves competitive overall conditional distribution recovery, and preserves mixed-type dependence structure on synthetic and real-world datasets.

By Siyuan Tang, Gongjun Xu, Ji Zhu
arXiv AI
Aug 24

GRALIS: Fusing Coalition and Gradient Attribution with Closed-Form Conservation Error and Finite-Sample Guarantees

GRALIS (Gradient‑Riesz Averaged Locally‑Integrated Shapley) merges coalition‑based and gradient‑based post‑hoc XAI techniques into a single estimator. It provides two certified guarantees: an exact closed‑form completeness deficit and a finite‑sample bound on the self‑normalized ratio. The method is grounded in a representation‑theoretic result that uniquely characterizes additive, linear, continuous attribution functionals, and it is experimentally illustrated on breast histology imaging.

By Raimondo Fanale
Hugging Face Trending Papers
Sep 17

Evaluating Explanation Methods by the Predictors They Induce

The paper introduces a straightforward evaluation method for explanation techniques: by converting each explanation into a predictor that sums the feature effects, the authors assess how accurately this predictor reproduces the original model’s predictions on unseen data. This approach applies to any explanation expressible as a function of features and is demonstrated on PDP, ALE, SHAP, and LIME. The authors theoretically show that summing partial dependence curves yields the optimal additive summary when features are independent, but this property fails with dependent features, and empirical results across diverse datasets confirm that the best-performing method depends on feature dependence.