An Analysis of Residual-Stream Geometry Across Transformer Depth
arXiv:2607. 18348v1 Announce Type: cross Abstract: We propose a transition-centred geometric analysis of transformer residual streams.
arXiv:2607. 18348v1 Announce Type: cross Abstract: We propose a transition-centred geometric analysis of transformer residual streams.
arXiv:2608. 12447v1 Announce Type: new Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream.
The paper introduces effective depth (Deff), a scalar diagnostic that treats a transformer’s layer‑wise residual stream as a discrete‑time process and measures how representation similarity decays with layer distance. Across sixteen decoder‑only language models, Deff reveals that most models exhibit a lower similarity decay than the closed‑form reference, indicating correlated residual updates rather than unused depth. The study also shows that this effect is robust to various controls and persists early in training, suggesting Deff is a global accumulated‑state diagnostic rather than a capability score.
arXiv:2606. 07559v2 Announce Type: replace-cross Abstract: Fine-tuning a language model often fails silently when its correct completion must outrank a near-synonym competitor.
arXiv:2603.18908v5 Announce Type: replace Abstract: Independently trained language models often learn compatible late-stage representations, despite differences in training objectives, architectures,...
arXiv:2606. 07559v1 Announce Type: cross Abstract: Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently.
arXiv:2609.15975v1 Announce Type: cross Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...
arXiv:2606. 02765v1 Announce Type: cross Abstract: Model dimension ($d_{model}$) is a fundamental hyperparameter in transformer language models, yet its role in setting the geometric limits of feature representation remains under-explored.
arXiv:2609.37717v1 Announce Type: new Abstract: Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state thro...
arXiv:2609.24209v1 Announce Type: new Abstract: The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent wor...
arXiv:2603. 06592v2 Announce Type: replace-cross Abstract: Contemporary studies in mechanistic interpretability have uncovered many puzzling phenomena in the neural information processing of Transformer-based language models, such as induction heads, function vectors, and the Hydra effect.
arXiv:2607. 10578v1 Announce Type: new Abstract: Existing hypotheses represent a concept in an LLM as a single point, a linear direction, or a Gaussian cluster, yet it remains unclear how and why such structures emerge.