The paper introduces a flexible symbolic framework that efficiently computes logical explanations for deep neural networks by parameterizing explanations with internal neuron activations and leveraging general-purpose logical engines like SMT solvers. Unlike previous methods that rely on specialized verifiers or are limited to individual input features, this approach is not restricted in shape and can scale to deep architectures. Experiments on image recognition and medical benchmarks demonstrate improved computational efficiency and the ability to explain networks that were previously intractable for logic-based methods.
By Tom\'a\v{s} Kol\'arik, Faezeh Labbaf, Fabrizio Leopardi, Grigory Fedyukovich, Michael Wand, Natasha Sharygina
arXiv:2608. 03772v1 Announce Type: new Abstract: Explaining the predictions of neural networks is a central challenge in trustworthy AI.
By Jannick Strobel, Muqsit Azeem, Stefan Leue
arXiv:2607. 07316v1 Announce Type: new Abstract: This article offers a comprehensive overview of mechanistic interpretability, an emerging field that seeks to reverse-engineer the internal algorithms of modern neural networks.
By Pranav Sawant, Jakub Krej\v{c}\'i
arXiv:2504.14015v2 Announce Type: replace-cross
Abstract: We introduce "causal pieces", a novel concept for analysing spiking neural networks (SNNs), inspired by "linear pieces" used to study express...
By Dominik Dold, Philipp Christian Petersen
arXiv:2411. 08875v4 Announce Type: replace Abstract: Existing algorithms for explaining the output of image classifiers use different definitions of explanations and a variety of techniques to find them.
By Hana Chockler, David A. Kelly, Daniel Kroening, Youcheng Sun
arXiv:2609.06862v1 Announce Type: new
Abstract: Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons...
By Dai Shi, Xiaoyu Li, Andi Han, Jos\'e Miguel Hern\'andez-Lobato
arXiv:2508. 11214v2 Announce Type: replace-cross Abstract: Explanations of cognitive behavior often appeal to computations over representations.
By Atticus Geiger, Jacqueline Harding, Thomas Icard
arXiv:2510. 14538v3 Announce Type: replace Abstract: Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.
By Emanuele Marconato, Samuele Bortolotti, Emile van Krieken, Paolo Morettin, Elena Umili, Antonio Vergari, Efthymia Tsamoura, Andrea Passerini, Stefano Teso
arXiv:2510. 04500v3 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves performance.
By Linghao Kong, Inimai Subramanian, Yonadav Shavit, Micah Adler, Dan Alistarh, Nir Shavit
The paper investigates the expressivity of time-to-first-spike spiking neural networks, showing that each neuron's firing time can be represented in a maxout-like form with many constrained affine pieces. It formalizes causal regions as polyhedral sets defined by fixed causal spike sequences and derives bounds on the number of such regions for both shallow and multilayer networks. Experiments confirm that spiking networks can produce richer input-space partitions than conventional feedforward ReLU networks.
By Manjot Singh, Guido Mont\'ufar, Gitta Kutyniok
arXiv:2608. 06839v1 Announce Type: new Abstract: Artificial Neural networks (ANNs) are often treated as black-box models, making explainability a central challenge in deep learning.
By Quanshi Zhang, Qihan Ren, Siyu Lou
arXiv:2603. 23867v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have been applied to a wide range of reasoning tasks, yet it remains unclear whether they can reason robustly under distribution shifts.
By Weixin Chen, Antonio Vergari, Han Zhao