arXiv:2508.10148v2 Announce Type: replace-cross
Abstract: Accurate and explainable out-of-distribution (OOD) detection is required to use machine learning systems safely. Previous work has shown that...
By Maria Stoica, Francesco Leofante, Alessio Lomuscio
The paper introduces a hypothesis‑testing framework that embeds feature importance methods (FIMs) within a Weight of Evidence (WoE) analysis. By quantifying how strongly observed evidence supports a given hypothesis—whether from domain knowledge, ground truth, or the FIM itself—the approach evaluates FIM alignment and variability. The authors provide theoretical links between WoE and attribution variance and demonstrate the method on LIME and SHAP explanations across varied reference hypotheses.
By Eddie Conti, Claudio Daka, \'Alvaro Parafita, Antonio L. Alfeo, Axel Brando, Mario G. C. A. Cimino
arXiv:2608. 02697v1 Announce Type: cross Abstract: Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models.
By Eddie Conti, \'Alvaro Parafita, Axel Brando
arXiv:2604. 15107v2 Announce Type: replace-cross Abstract: Shapley values provide a flexible framework for attributing feature contributions to model predictions, but they are not naturally suited for feature selection: a feature may receive a positive attribution even when it is redundant given the remaining variables.
By Chenghui Zheng, Garvesh Raskutti
arXiv:2607. 24145v1 Announce Type: new Abstract: Feature selection aims to identify the most informative and relevant features for a given dataset, either in terms of capturing the underlying data structure and distribution better, or with respect to the performance on a downstream task.
By Muhammad Rajabinasab, Arthur Zimek
The paper introduces a unified framework called null importance to clarify different notions of feature relevance in interpretable machine learning. It defines null importance at the population level for various relevance concepts—marginal, conditional, predictive risk, functional invariance, and causal effects—and demonstrates how each answers distinct scientific questions. Through theoretical analysis, simulations, and case studies on fairness and genomic modeling, the authors show when these null notions coincide or diverge and how different importance methods target them.
By Garvesh Raskutti, Kris Sankaran, Jiaxin Ye
The paper proposes a new evaluation test for explanation methods: if an explanation accurately captures how a model uses its features, one should be able to reconstruct the model’s predictions from it. The authors convert explanations into predictors by summing feature effects and assess how well these predictors reproduce the model on unseen data, without any fitting. They apply this test to partial dependence plots, accumulated local effects, SHAP, and LIME across multiple datasets and model families, showing that the best method depends on feature dependence and that some existing quality metrics can favor flawed explanations.
By Jacob Selb{\ae}k, Hugo L. Hammer
The paper introduces a straightforward evaluation method for explanation techniques: by converting each explanation into a predictor that sums the feature effects, the authors assess how accurately this predictor reproduces the original model’s predictions on unseen data. This approach applies to any explanation expressible as a function of features and is demonstrated on PDP, ALE, SHAP, and LIME. The authors theoretically show that summing partial dependence curves yields the optimal additive summary when features are independent, but this property fails with dependent features, and empirical results across diverse datasets confirm that the best-performing method depends on feature dependence.
arXiv:2607. 16478v1 Announce Type: cross Abstract: Oversampling is widely used to address class imbalance in tabular classification, but existing methods can distort the feature importance ranking underlying model explanations.
By Marios Tyrovolas, Argiris Sofotasios, Dimitris Metaxakis, Georgios Mermigkis, George Georgoulas, Panagiotis Hadjidoukas, Chrysostomos Stylios
arXiv:2607. 15774v1 Announce Type: cross Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored.
By Davide Italo Serramazza, Thach Le Nguyen, Georgiana Ifrim
arXiv:2607. 22045v1 Announce Type: new Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome.
By Oleksii Furman, {\L}ukasz Lenkiewicz, Marcel Musia{\l}ek, Maciej Zi\k{e}ba
ProToMEx is a new explainability framework that uses Probabilistic Topic Models to learn latent topics representing high‑level reasons behind a classifier’s decisions, moving beyond simple feature attribution. It provides both global and local explanations, revealing multiple co‑existing reasons for individual predictions. Empirical results show that ProToMEx achieves comparable fidelity to SHAP and LIME while being 30–40× faster on standard tabular and synthetic datasets.
By Athina Georgara, Adarsh Valoor, Sarvapali D. Ramchurn