arXiv Machine Learning

When Confidence Lacks Concepts: Interpretable OOD Detection via Representation Perturbations

arXiv:2606. 16196v1 Announce Type: new Abstract: Deep neural networks have achieved remarkable performance across medical imaging tasks, yet their tendency to overgeneralize under distributional shifts poses a major obstacle to safe clinical deployment.

arXiv Computer Vision
Sep 7

FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders

The paper introduces FailSAE, a method that uses Sparse Autoencoders to predict failures in vision‑language models (VLMs) such as CLIP. By treating failure prediction as a classification over sparse SAE latent activations and employing a three‑stage training pipeline, the approach yields higher prediction accuracy than existing confidence‑score or auxiliary‑classifier baselines. Analysis shows that the SAE captures class‑specific concepts and reveals a shift toward ambiguous or style‑related concepts during failures, offering insights for runtime failure recovery.

By Jie Ma, Zongxi Liu, Yi Zhu
arXiv Machine Learning
Aug 31

EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders

EXPOSE is a framework that applies Sparse Autoencoders to Vision Foundation Model embeddings in computational pathology, aiming to separate biological signals from domain‑specific noise. By training a sparse representation of VFM features and using a linear classifier to flag domain‑specific latent dimensions, the method masks these components before downstream relapse prediction, avoiding the need to retrain the backbone model. Experiments on a large prostate cancer dataset demonstrate that removing domain‑specific features improves cross‑domain performance and raises the Domain Robustness Index (DoRI).

By Anja Witte, Maximilian Lennartz, Jan Baumbach, Guido Sauter, Stefan Bonn, Patrick Fuhlert, Marina Zimmermann
arXiv Machine Learning
Aug 27

ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

ICON Decomposition is a new method for explaining deep neural networks by quantifying how much variance each concept explains in a network layer after accounting for all other concepts and the outcome. Unlike previous concept‑based methods that evaluate concepts in isolation, ICON can distinguish genuine model reliance from spurious correlations. Experiments on synthetic data, skin‑lesion, and brain‑imaging models show that ICON recovers concept importance more accurately, isolates truly relied‑upon concepts, and provides sparse explanations validated through retraining and out‑of‑distribution testing.

By Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, Sonja Greven, Kerstin Ritter
arXiv AI
Jun 24

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

arXiv:2606. 24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence.

By Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir