arXiv AI

ICON Decomposition: Auditing deep neural networks for shortcuts by decomposing layer-wise representations using concepts

The paper introduces ICON Decomposition, a method for auditing deep neural networks by decomposing layer-wise representations into independent concept contributions. Unlike existing techniques that rely on linear probes or concept activation vectors, ICON quantifies the variance share each concept explains while conditioning on all other concepts and the outcome, allowing comparison across layers and concept types. Experiments on simulated data, skin‑cancer, and neuroimaging models show that ICON more accurately recovers true concept importance and can distinguish learned shortcuts from correlated concepts, as validated by retraining and out‑of‑distribution tests.

arXiv AI
Sep 1

ICON Decomposition: Auditing Deep Neural Networks with Multivariate Variance-based Concept-level Explanations

arXiv:2608.26083v2 Announce Type: replace-cross Abstract: Deep neural networks often exploit spurious associations, a failure known as shortcut learning. Auditing for shortcuts requires testing many...

By Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, Sonja Greven, Kerstin Ritter
arXiv Machine Learning
Aug 27

ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

ICON Decomposition is a new method for explaining deep neural networks by quantifying how much variance each concept explains in a network layer after accounting for all other concepts and the outcome. Unlike previous concept‑based methods that evaluate concepts in isolation, ICON can distinguish genuine model reliance from spurious correlations. Experiments on synthetic data, skin‑lesion, and brain‑imaging models show that ICON recovers concept importance more accurately, isolates truly relied‑upon concepts, and provides sparse explanations validated through retraining and out‑of‑distribution testing.

By Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, Sonja Greven, Kerstin Ritter
arXiv AI
Jul 7

TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models

arXiv:2607. 04593v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed to the language model.

By Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda, Alaa Eddine Mazouz, Van-Tam Nguyen
arXiv AI
Aug 12

Token-Based Detection of Spurious Correlations in Vision Transformers

arXiv:2509. 04009v2 Announce Type: replace-cross Abstract: Due to their powerful feature association capabilities, neural network-based computer vision models have the ability to detect and exploit unintended patterns within the data, potentially leading to correct predictions based on incorrect or unintended but statistically relevant signals.

By Solha Kang, Esla Timothy Anzaku, Wesley De Neve, Arnout Van Messem, Joris Vankerschaver, Francois Rameau, Utku Ozbulak