Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Prob...
arXiv:2608. 12447v1 Announce Type: new Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream.
The paper introduces Conditional Functional Substitutability (CFS) as a new way to measure redundancy in Transformers by examining when intermediate states produce similar downstream responses. CFS uncovers functional relationships and potential reductions that traditional importance- or similarity-based metrics miss, revealing systematic reorganization as models scale. Experiments across modalities and Transformer families show that performance gains do not always align with increased substitutability, and that models with more independent functional structure perform better, offering a functional explanation for diminishing returns and enabling more efficient dynamic computation.
arXiv:2608. 01968v1 Announce Type: new Abstract: Transformer models are most often understood through what they do: their benchmark performance, generation quality, or behavior on downstream tasks.
arXiv:2609.16537v1 Announce Type: cross Abstract: Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical kn...
arXiv:2605. 27458v2 Announce Type: replace-cross Abstract: Transformer has significantly propelled the development of artificial intelligence, and certainly the development of agents as well.