arXiv AI

The Misery of Mechanistic Interpretability: A Formal Perspective

arXiv Machine Learning
Jul 1

Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

arXiv:2606. 32008v1 Announce Type: new Abstract: Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens.

By Philippe Chlenski, Zachariah Carmichael, Ayush Warikoo, Chia-Tse Shao, Yingxiao Ye, Aobo Yang, Vivek Miglani, Nehal Bandi
arXiv AI
Aug 11

Scaling Inherently Interpretable Language Models

arXiv:2608. 07594v1 Announce Type: cross Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish.

By Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail, Giang Nguyen, Isaac Plant, Muawiz Chaudhary, Nathaniel Monson, Saqib Azim, Zhichen Guo, Julius Adebayo
arXiv Machine Learning
Jun 18

From Mechanistic to Compositional Interpretability

arXiv:2605. 08934v2 Announce Type: replace Abstract: Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable components.

By Ward Gauderis, Thomas Dooms, Steven T. Homer, Kola Ayonrinde, Geraint A. Wiggins
arXiv Machine Learning
Sep 11

Beyond Solver Verdicts: Generative Reward Models for Autoformalization

The paper identifies a new failure mode in neurosymbolic systems called Verdict‑Preserving‑Unfaithfulness (VPU), where incorrect formal encodings can still pass solver checks. It introduces Generative Verification (GenV), a method that uses a language model to produce a continuous reference‑equivalence score without relying on explicit localization. Experiments show GenV+HN achieves high AUROC, generalizes to unseen translators, and improves downstream agent performance by 11.3 points.

By Vikash Singh, Debargha Ganguly, Aman Goel, Ali Torkamani, Xiaoxue Han, Joseph Lilien, Ferhat Erata, Vipin Chaudhary
arXiv AI
3d ago

Generative Interpretability via Scalable Neuro-Symbolic Models

The paper argues that as Large Language Models transition from chatbots to agentic systems, the current post-hoc interpretability paradigm is insufficient for safe deployment because it cannot audit or intervene before an output is produced. It proposes a shift to generative interpretability, where a model’s inference process inherently exposes semantically meaningful checkpoints that are human-understandable and can be causally intervened upon. The authors illustrate the advantages of this approach and introduce Neuro‑Symbolic Models as a concrete implementation.

By Xiaocong Yang
arXiv Machine Learning
Jul 2

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

arXiv:2607. 01033v1 Announce Type: new Abstract: Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques.

By Andrzej Szablewski, Gabriel Konar-Steenberg, Raffaello Fornasiere, Nikita Menon, Stefan Heimersheim
arXiv AI
Aug 26

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

The paper introduces a formal auditing framework to evaluate the robustness and fidelity of post‑hoc explainers such as SHAP and LIME. It defines a Trust Score that combines how stable an explanation is under small input perturbations with how well the highlighted features actually influence the model’s prediction. Experiments on a Madagascar malnutrition dataset show that even highly accurate models can produce unreliable explanations, and that fidelity scores degrade when models overfit.

By Rosa Elysabeth Ralinirina, Jean Christian Ralaivao, Niaiko Micha\"el Ralaivao, Alain Josu\'e Ratovondrahona, Thomas Mahatody