arXiv AI By Daniel A. Herrmann, Benjamin A. Levinstein

Radical AI Interpretability

Read the original on arXiv AI →

arXiv:2606. 26523v1 Announce Type: new Abstract: We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 7

From Interpretability Methods to Interpretable Models

The paper argues that explainable AI for computer vision has focused too much on developing interpretability methods rather than assessing how interpretable the models themselves are. It proposes a shift toward model-centric evaluation, using existing tools to compare what different models represent and compute, and emphasizes the need to measure whether humans can truly understand these models. The authors review the current toolbox, survey limited model comparison work, draw parallels to systems neuroscience, and outline a future agenda for model-focused XAI.

By Julien Colin, Nuria Oliver, Thomas Serre
arXiv AI
Jun 11

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

arXiv:2606. 12289v1 Announce Type: cross Abstract: As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations.

By Pietro Barbiero, Giovanni De Felice, Mateo Espinosa Zarlenga, Francesco Giannini, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra, Ruggero Noris
arXiv AI
Sep 18

Xeno-Interpretability: Investigating the Alien Minds of LLMs

The paper introduces the concept of xeno-interpretability, which studies internal distinctions in large language models that lack corresponding human concepts. It distinguishes between human‑interpretable and xeno‑semantic spaces, showing that LLMs possess a far larger internal representational space than can be captured by finite human descriptions. The authors propose an empirical program to identify and characterize these xeno‑representations, noting their potential to influence model behavior in ways that are not fully visible through human‑readable communication.

By F. Pierucci, M. Bracale Syrnikov, M. Prandi, M. Galisai, F. Giarrusso, P. Bisconti
arXiv AI
Sep 15

Generative Interpretability via Scalable Neuro-Symbolic Models

The paper argues that as Large Language Models transition from chatbots to agentic systems, the current post-hoc interpretability paradigm is insufficient for safe deployment because it cannot audit or intervene before an output is produced. It proposes a shift to generative interpretability, where a model’s inference process inherently exposes semantically meaningful checkpoints that are human-understandable and can be causally intervened upon. The authors illustrate the advantages of this approach and introduce Neuro‑Symbolic Models as a concrete implementation.

By Xiaocong Yang
arXiv AI
Jun 15

Actionable Interpretability Must Be Defined in Terms of Symmetries

arXiv:2601. 12913v4 Announce Type: replace Abstract: This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpretability can be formally tested or designed for.

By Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini, Alberto Termine, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra
arXiv AI
Sep 2

Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective

The paper examines how large language model (LLM) outputs are increasingly used in contexts that demand justified interpretations, such as law, education, policy analysis, and public moral debate. It identifies a recurring failure—interpretive misplacement—where model-generated readings are treated as settled meanings without explicit interpretive frames, provenance, or defensible alternatives, leading to accountability loss. Drawing on philosophical hermeneutics, the author proposes design principles for human‑AI co‑interpretation, reorganizes existing LLM techniques into hermeneutically responsible patterns, and discusses implications for legal practice, education, scholarship, and public discourse, while framing digital hermeneutics as a literacy for critically engaging with AI‑mediated texts.

By Behrooz Razeghi