arXiv Machine Learning

Beyond $L_2$: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures

arXiv:2608. 16773v1 Announce Type: new Abstract: Prototype-based neural networks are hailed as interpretable-by-design architectures.

arXiv Machine Learning
Jun 2

Normalized Relevance Measure as a Unifying Framework to Explain Neural Network Latent Structures

arXiv:2606. 00557v1 Announce Type: new Abstract: To understand how a neural network (NN) functions and makes predictions, it has become increasingly clear that analyzing only the input domain is insufficient -- one must also examine its internal inference mechanisms to capture the complete picture.

By Ping Xiong, Thomas Schnake, Gr\'egoire Montavon, Klaus-Robert M\"uller, Shinichi Nakajima
arXiv Machine Learning
Sep 25

Pointwise Generalization in Deep Neural Networks

The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.

By Shaojie Li, Yunbei Xu
arXiv Machine Learning
Sep 15

Neuron Activation-based Computation of Logical Explanations for Deep Neural Networks

The paper introduces a flexible symbolic framework that efficiently computes logical explanations for deep neural networks by parameterizing explanations with internal neuron activations and leveraging general-purpose logical engines like SMT solvers. Unlike previous methods that rely on specialized verifiers or are limited to individual input features, this approach is not restricted in shape and can scale to deep architectures. Experiments on image recognition and medical benchmarks demonstrate improved computational efficiency and the ability to explain networks that were previously intractable for logic-based methods.

By Tom\'a\v{s} Kol\'arik, Faezeh Labbaf, Fabrizio Leopardi, Grigory Fedyukovich, Michael Wand, Natasha Sharygina
Hugging Face Trending Papers
Aug 19

Graphical Design of Interpretable Architectures

The paper introduces a new graphical notation, adapted from Penrose tensor notation, to design and represent interpretable AI architectures. Unlike symbolic equations or probabilistic models, this notation provides a global view of an architecture while directly mapping to PyTorch einsum code. The authors demonstrate its use on several interpretable models and on the Steerling-8B language model, revealing structural insights and enabling concise code generation.