arXiv Machine Learning

FMMO: Detecting the Divergence Between Local Attribution and Global Drift

arXiv AI
Aug 26

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

The paper introduces a formal auditing framework to evaluate the robustness and fidelity of post‑hoc explainers such as SHAP and LIME. It defines a Trust Score that combines how stable an explanation is under small input perturbations with how well the highlighted features actually influence the model’s prediction. Experiments on a Madagascar malnutrition dataset show that even highly accurate models can produce unreliable explanations, and that fidelity scores degrade when models overfit.

By Rosa Elysabeth Ralinirina, Jean Christian Ralaivao, Niaiko Micha\"el Ralaivao, Alain Josu\'e Ratovondrahona, Thomas Mahatody
arXiv AI
Aug 28

Diagnosing Conformal Prediction Failures Under Distribution Shift: A COVID-19 Case Study

The paper introduces SHAP concentration as a pre‑deployment diagnostic for detecting when conformal prediction will fail under distribution shift, specifically in gradient‑boosted classifiers. Using a COVID‑19 supply‑chain case study, the authors show that higher feature‑importance concentration correlates with larger drops in coverage, while standard shift detectors cannot differentiate between catastrophic and robust outcomes. The diagnostic is validated on additional datasets, and a formal theorem links concentration to worsening conformity‑score bounds, though it does not capture global‑sensitivity failures in neural networks.

By Chorok Lee
arXiv AI
4d ago

Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations

The paper introduces a unified evaluation protocol for robust counterfactual explanations (CFE), testing six robust methods and two baselines across four tabular datasets under eight types of model change. It shows that robustness scores vary by change type and that methods designed for one change family may not transfer to others, with RobX performing most consistently. The study emphasizes the need for a common protocol that defines model changes, measures their impact, and separates generation performance from robustness.

By Marcin Kostrzewa, Maciej Zi\k{e}ba
Hugging Face Trending Papers
Jul 2

Beyond Gradient-Based Attacks: Adversarial Robustness and Explainability Stability in Cybersecurity Classifiers

Adversarial attacks on cybersecurity classifiers pose a dual threat: degrading predictions and destabilising the SHAP-based explanations that security analysts rely on to understand and triage alerts. We extend our prior MLP conference study to Random Forest and XGBoost across four tabular security datasets (phishing URLs, UNSW-NB15, NF-ToN-IoT, HIKARI-2021), evaluating five attacks including three black-box methods applicable to non-differentiable tree models.

arXiv AI
Jun 9

REFLECT: Intervention-Supported Error Attribution for Silent Failures in LLM Agent Traces

arXiv:2606. 09071v1 Announce Type: new Abstract: Large language model (LLM) agents now solve complex tasks through long plan-and-execution traces, yet the ability to locate errors in a completed traces still lags far behind, especially in the \emph{silent failure} regime.

By Xiaofeng Lin, Yingxu Wang, Tung Sum Thomas Kwok, Daniel Guo, Sahil Arun Nale, Charles Fleming, Guang Cheng