arXiv Machine Learning By Tiejin Chen, Wenwang Huang, Linsey Pang, Dongsheng Luo, Hua Wei

Are Classification Robustness and Explanation Robustness Really Strongly Correlated? An Analysis Through Input Loss Landscape

Read the original on arXiv Machine Learning →

arXiv:2403. 06013v2 Announce Type: replace Abstract: This paper delves into the critical area of deep learning robustness, challenging the conventional belief that classification robustness and explanation robustness in image classification systems are inherently correlated.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 23

eXplaining to Learn (eX2L): Regularization Using Contrastive Visual Explanation Pairs for Distribution Shifts

The paper introduces eXplaining to Learn (eX2L), an interpretable framework that regularizes a classifier by penalizing similarity between Grad‑CAM maps of the main label classifier and a confounder classifier. This approach decorrelates confounding features from latent representations during training. On the Spawrious Many‑to‑Many Hard Challenge benchmark, eX2L outperforms the current state‑of‑the‑art by 5.49% in average accuracy and 10.90% in worst‑group accuracy, while also demonstrating functional domain invariance through explicit label‑nuisance decoupling.

By Paulo Mario P. Medina, Jose Marie Antonio Mi\~noza, Sebastian C. Iba\~nez
arXiv Machine Learning
Aug 27

ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

ICON Decomposition is a new method for explaining deep neural networks by quantifying how much variance each concept explains in a network layer after accounting for all other concepts and the outcome. Unlike previous concept‑based methods that evaluate concepts in isolation, ICON can distinguish genuine model reliance from spurious correlations. Experiments on synthetic data, skin‑lesion, and brain‑imaging models show that ICON recovers concept importance more accurately, isolates truly relied‑upon concepts, and provides sparse explanations validated through retraining and out‑of‑distribution testing.

By Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, Sonja Greven, Kerstin Ritter