arXiv Machine Learning By Umberto Biccari, Qian Huang, Enrique Zuazua

Interior interpretability with attention rollout: contraction and propagation profiles in Transformers

Read the original on arXiv Machine Learning →

arXiv:2607. 22367v1 Announce Type: new Abstract: Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 7

Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

The paper introduces an influence score that measures how much each attention head contributes to classification decisions in Transformer models, specifically for prompt injection detection. The score blends directional effects on logits with structural impact within the residual stream, allowing analysis at head, layer, and network scales. When applied to a DeBERTa model, the framework uncovers different decision patterns for correct versus incorrect predictions, offering a balanced approach between detailed circuit analysis and global output methods.

By Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi
arXiv AI
Jun 2

A Monosemantic Attribution Framework for Stable Interpretability in Clinical Neuroscience Transformer-Based Language Models

arXiv:2601. 17952v2 Announce Type: replace-cross Abstract: Interpretability remains a key challenge for deploying language models (LM) in clinical settings such as progression diagnosis of Alzheimer disease, where early and trustworthy predictions are essential.

By Michail Mamalakis, Tiago Azevedo, Cristian Cosentino, Chiara D'Ercoli, Subati Abulikemu, Zhongtian Sun, Richard Bethlehem, Pietro Lio
arXiv AI
Sep 25

Every Component Is a Lookup: One Linear Graph for Interaction, Composition and Attribution

The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.

By Po-Kai Chen, Aske Plaat, Niki van Stein