arXiv AI

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

arXiv:2606. 12289v1 Announce Type: cross Abstract: As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations.

arXiv AI
Jun 26

Radical AI Interpretability

arXiv:2606. 26523v1 Announce Type: new Abstract: We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability.

By Daniel A. Herrmann, Benjamin A. Levinstein
arXiv Machine Learning
Jun 18

From Mechanistic to Compositional Interpretability

arXiv:2605. 08934v2 Announce Type: replace Abstract: Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable components.

By Ward Gauderis, Thomas Dooms, Steven T. Homer, Kola Ayonrinde, Geraint A. Wiggins
arXiv Computer Vision
Sep 7

From Interpretability Methods to Interpretable Models

The paper argues that explainable AI for computer vision has focused too much on developing interpretability methods rather than assessing how interpretable the models themselves are. It proposes a shift toward model-centric evaluation, using existing tools to compare what different models represent and compute, and emphasizes the need to measure whether humans can truly understand these models. The authors review the current toolbox, survey limited model comparison work, draw parallels to systems neuroscience, and outline a future agenda for model-focused XAI.

By Julien Colin, Nuria Oliver, Thomas Serre
arXiv AI
Jun 15

Actionable Interpretability Must Be Defined in Terms of Symmetries

arXiv:2601. 12913v4 Announce Type: replace Abstract: This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpretability can be formally tested or designed for.

By Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini, Alberto Termine, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra
arXiv Machine Learning
Jul 2

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

arXiv:2607. 01033v1 Announce Type: new Abstract: Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques.

By Andrzej Szablewski, Gabriel Konar-Steenberg, Raffaello Fornasiere, Nikita Menon, Stefan Heimersheim
arXiv AI
Sep 15

Data storytelling meets interpretable machine learning: Decoding AI decisions for non-experts without revealing sensitive data and model details

The paper proposes a framework that blends data storytelling with interpretable machine learning (IML) to make AI decisions more understandable to non-experts while protecting sensitive data. It introduces formal concepts such as the DIST Pyramid and the I-P-O Model, and builds an architecture that generates "What-if" and "Why-not" explanations using SHAP values and large language models. A case study on the Boston Housing dataset shows that participants found these data stories significantly more comprehensible and accessible than traditional SHAP visualizations.

By Lemen Chao, Zixuan Yang, Anran Fang, Mingran Sun, Ming Lei
arXiv AI
Jul 17

Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models

arXiv:2607. 14315v1 Announce Type: cross Abstract: In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across multiple datasets and machine learning models, with the ultimate goal of creating a unified multidimensional explainability score.

By Georgios Makridis, Georgios Fatouros, Athanasios Kiourtis, Dimitrios Kotios, Vasileios Koukos, Dimosthenis Kyriazis, Jonh Soldatos
arXiv AI
Sep 10

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

The paper introduces SAEScientist-Bench, a benchmark that tests whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders (SAEs). Agents are tasked with designing contrastive probes and navigating a large feature dictionary in Gemma-2-9B-IT to identify optimal features for a target concept, with performance measured against expert-curated references on activation rank, concept selectivity, and causal steering. Results show that while frontier agents can discover features and outperform controls, they still lag behind expert baselines, especially in causal steering, highlighting both the potential and current limitations of closed-loop autonomous AI research.

By Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu
arXiv AI
Jul 16

AIMO Interpretability Challenge

arXiv:2607. 13899v1 Announce Type: new Abstract: We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models' internal mechanisms.

By Michal \v{S}tef\'anik, Philipp Mondorf, Andreas Waldis, Qianying Liu, Chuan Yang, Michal Spiegel, Josef Kucha\v{r}, Marek Kadl\v{c}\'ik, Adam Vawda-Oomerjee, Chaoran Liu, Simon Frieder, Barbara Plank, Fazl Barez, Pontus Stenetorp
arXiv AI
Sep 15

Generative Interpretability via Scalable Neuro-Symbolic Models

The paper argues that as Large Language Models transition from chatbots to agentic systems, the current post-hoc interpretability paradigm is insufficient for safe deployment because it cannot audit or intervene before an output is produced. It proposes a shift to generative interpretability, where a model’s inference process inherently exposes semantically meaningful checkpoints that are human-understandable and can be causally intervened upon. The authors illustrate the advantages of this approach and introduce Neuro‑Symbolic Models as a concrete implementation.

By Xiaocong Yang