The paper introduces a flexible symbolic framework that efficiently computes logical explanations for deep neural networks by parameterizing explanations with internal neuron activations and leveraging general-purpose logical engines like SMT solvers. Unlike previous methods that rely on specialized verifiers or are limited to individual input features, this approach is not restricted in shape and can scale to deep architectures. Experiments on image recognition and medical benchmarks demonstrate improved computational efficiency and the ability to explain networks that were previously intractable for logic-based methods.
By Tom\'a\v{s} Kol\'arik, Faezeh Labbaf, Fabrizio Leopardi, Grigory Fedyukovich, Michael Wand, Natasha Sharygina
arXiv:2608. 03772v1 Announce Type: new Abstract: Explaining the predictions of neural networks is a central challenge in trustworthy AI.
By Jannick Strobel, Muqsit Azeem, Stefan Leue
arXiv:2607. 07316v1 Announce Type: new Abstract: This article offers a comprehensive overview of mechanistic interpretability, an emerging field that seeks to reverse-engineer the internal algorithms of modern neural networks.
By Pranav Sawant, Jakub Krej\v{c}\'i
arXiv:2504.14015v2 Announce Type: replace-cross
Abstract: We introduce "causal pieces", a novel concept for analysing spiking neural networks (SNNs), inspired by "linear pieces" used to study express...
By Dominik Dold, Philipp Christian Petersen
arXiv:2411. 08875v4 Announce Type: replace Abstract: Existing algorithms for explaining the output of image classifiers use different definitions of explanations and a variety of techniques to find them.
By Hana Chockler, David A. Kelly, Daniel Kroening, Youcheng Sun
arXiv:2609.06862v1 Announce Type: new
Abstract: Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons...
By Dai Shi, Xiaoyu Li, Andi Han, Jos\'e Miguel Hern\'andez-Lobato