arXiv Machine Learning

Are Classification Robustness and Explanation Robustness Really Strongly Correlated? An Analysis Through Input Loss Landscape

arXiv:2403. 06013v2 Announce Type: replace Abstract: This paper delves into the critical area of deep learning robustness, challenging the conventional belief that classification robustness and explanation robustness in image classification systems are inherently correlated.

arXiv Machine Learning
Sep 23

eXplaining to Learn (eX2L): Regularization Using Contrastive Visual Explanation Pairs for Distribution Shifts

The paper introduces eXplaining to Learn (eX2L), an interpretable framework that regularizes a classifier by penalizing similarity between Grad‑CAM maps of the main label classifier and a confounder classifier. This approach decorrelates confounding features from latent representations during training. On the Spawrious Many‑to‑Many Hard Challenge benchmark, eX2L outperforms the current state‑of‑the‑art by 5.49% in average accuracy and 10.90% in worst‑group accuracy, while also demonstrating functional domain invariance through explicit label‑nuisance decoupling.

By Paulo Mario P. Medina, Jose Marie Antonio Mi\~noza, Sebastian C. Iba\~nez
arXiv Machine Learning
Aug 27

ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

ICON Decomposition is a new method for explaining deep neural networks by quantifying how much variance each concept explains in a network layer after accounting for all other concepts and the outcome. Unlike previous concept‑based methods that evaluate concepts in isolation, ICON can distinguish genuine model reliance from spurious correlations. Experiments on synthetic data, skin‑lesion, and brain‑imaging models show that ICON recovers concept importance more accurately, isolates truly relied‑upon concepts, and provides sparse explanations validated through retraining and out‑of‑distribution testing.

By Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, Sonja Greven, Kerstin Ritter
arXiv AI
Aug 24

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization proposes a new framework that improves deepfake detection and interpretability. The approach introduces Feature-robust Augmentation—diversified degradation-aware strategies combined with supervised contrastive learning and a mean-teacher architecture—to maintain accuracy on low-quality images. For explanations, it employs evidence-grounded preference optimization, guiding the model to focus on genuine manipulation traces by learning from chosen-rejected explanation pairs that omit evidence or inject irrelevant details. The method achieved first place in the ACM Multimedia 2026 Explainable Deepfake Detection Challenge and is publicly available on GitHub.

By Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu
arXiv AI
Sep 24

Do Center Biases Propagate? Robustness of Pathology Foundation Models in Whole-Slide Image Classification

The study investigates whether pathology foundation models (PFMs) carry center-related biases into whole-slide image (WSI) classification. By training models with increasing class-center correlations and evaluating six PFMs across four datasets and two MIL aggregators, the authors introduce the Area Under the Cramér's V Curve (AUCC) to measure both accuracy and degradation due to spurious correlations. Results reveal that center information propagates to WSI predictions, with robustness varying by PFM and MIL strategy, and that ComBat harmonization does not consistently improve robustness.

By Il\'an Carretero, Pablo Meseguer, Roc\'io del Amor, Valery Naranjo
arXiv AI
Sep 1

ICON Decomposition: Auditing Deep Neural Networks with Multivariate Variance-based Concept-level Explanations

arXiv:2608.26083v2 Announce Type: replace-cross Abstract: Deep neural networks often exploit spurious associations, a failure known as shortcut learning. Auditing for shortcuts requires testing many...

By Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, Sonja Greven, Kerstin Ritter