arXiv Machine Learning

Enhancing Visual Feature Attribution via Weighted Integrated Gradients

arXiv:2505. 03201v4 Announce Type: replace-cross Abstract: Integrated Gradients (IG) is a widely used attribution method in explainable AI, particularly in computer vision applications where reliable feature attribution is essential.

arXiv AI
Sep 16

ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers

The paper introduces ResLRP, an extension of Layer-wise Relevance Propagation that explicitly handles residual connections in Vision Transformers to prevent attribution explosion. It demonstrates that residual cancellation causes instability in ViT explanations, and that ResLRP improves faithfulness and localization across a wide range of ViT architectures, including Vision Language Models. The method also provides a diagnostic measure for predicting attribution degradation and successfully localizes Sparse Autoencoder features.

By Jim Berend, Reduan Achtibat, Daniel Sch\"affer, Alexander Binder, Wojciech Samek, Sebastian Lapuschkin, Maximilian Dreyer
arXiv Machine Learning
Jun 10

XtrAIn: Training-Guided Occlusion for Feature Attribution

arXiv:2606. 10877v1 Announce Type: new Abstract: Occlusion-based attribution methods provide an intuitive way to estimate feature importance by perturbing input features and measuring the resulting change in model output.

By Thodoris Lymperopoulos, Ioannis Kakogeorgiou, Denia Kanellopoulou
arXiv Machine Learning
Jun 9

Analysis of Information Theory for Explainable AI

arXiv:2507. 09092v2 Announce Type: replace-cross Abstract: With the intervention of machine vision in our crucial day to day necessities including healthcare and automated power plants, attention has been drawn to the internal mechanisms of convolutional neural networks, and the reason why the network provides specific inferences.

By Ram S Iyer