Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So
arXiv:2608. 10251v1 Announce Type: cross Abstract: A transformer's answer lives on one axis: the direction its unembedding reads.
The paper investigates how many transformer components influence a token prediction by measuring the absolute contribution of each unit and channel to the logit. It finds that thousands of components contribute to a single prediction, yet a small subset—often just dozens—carries the majority of the predictive mass. Across models ranging from 124 M to 7 B parameters, the proportion of the model involved in a prediction remains around one to three percent, independent of size, and the study demonstrates that specific components can be directly read and written to modify model behavior without additional training.
arXiv:2608. 10251v1 Announce Type: cross Abstract: A transformer's answer lives on one axis: the direction its unembedding reads.
arXiv:2605. 23393v2 Announce Type: replace-cross Abstract: Mechanistic interpretability of transformers requires identifying not just which components matter but how they compose into the computational route that produced a prediction.
arXiv:2607. 08946v1 Announce Type: new Abstract: A transformer can be built from operators that are legible by construction -- bounded, named units that read as fuzzy set operations rather than dense activations -- but legibility must be pressed for during training, and the pressure has a failure mode.
The paper introduces the Communication Map, a method that charts every potential communication channel in a transformer model using only its weights. It generalizes previous coupling metrics into a single coefficient covering all 18 connection classes, revealing that 70‑89% of head pairs are non‑randomly oriented and identifying strong or avoiding couplings. The authors demonstrate the map’s utility by recovering known induction circuits and uncovering a two‑dimensional stream subspace whose removal eliminates induction capabilities across several models.
The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
arXiv:2603.18908v5 Announce Type: replace Abstract: Independently trained language models often learn compatible late-stage representations, despite differences in training objectives, architectures,...
arXiv:2607. 02964v1 Announce Type: cross Abstract: A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does.
arXiv:2607. 23054v1 Announce Type: cross Abstract: Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achieving 81% KV-cache reduction during inference.
arXiv:2609.16537v1 Announce Type: cross Abstract: Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical kn...
arXiv:2607. 04319v1 Announce Type: cross Abstract: A companion paper showed that a transformer's feed-forward layer can be rebuilt from explicit fuzzy set operations - intersection, set-difference, and a self-forgetting sequence quantifier - so its hidden units read as named logical operators at no cost to language-model quality.
arXiv:2606. 07559v1 Announce Type: cross Abstract: Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently.