Feed-Forward Steering in Transformer Residual Dynamics
arXiv:2608. 02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere.
arXiv:2605. 25225v2 Announce Type: replace-cross Abstract: Mechanistic interpretability often studies Transformer behavior by intervening on internal activations through activation patching, causal tracing, path patching, and steering directions.
arXiv:2608. 02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere.
arXiv:2607. 17696v1 Announce Type: cross Abstract: We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions.
arXiv:2608.30720v1 Announce Type: new Abstract: Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied...
arXiv:2609.15975v1 Announce Type: cross Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...
arXiv:2608.22034v1 Announce Type: new Abstract: Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks...
arXiv:2609.37717v1 Announce Type: new Abstract: Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state thro...
The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.
arXiv:2602. 06883v3 Announce Type: replace Abstract: The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversarial robustness.
arXiv:2603.19742v3 Announce Type: replace-cross Abstract: Understanding the internal mechanisms of transformer-based large language models (LLMs) is crucial for their reliable deployment and effectiv...
arXiv:2607. 22367v1 Announce Type: new Abstract: Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers.
arXiv:2609.16537v1 Announce Type: cross Abstract: Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical kn...
The paper investigates why adaptive optimizers like Adam outperform SGD when fine‑tuning Transformers. It introduces gradient heterogeneity—the variation in gradient norms across parameter blocks—and shows, both theoretically and experimentally, that this heterogeneity, together with Hessian heterogeneity, hampers SGD convergence while sign‑based methods such as SignSGD are less affected. The study links the source of gradient heterogeneity to layer‑normalization placement, finding that Post‑LN architectures exhibit the strongest effect, and uses SignSGD as a tractable proxy to analyze Adam‑like behavior and learning‑rate scaling.