Feed-Forward Steering in Transformer Residual Dynamics
arXiv:2608. 02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere.
arXiv:2605. 25225v2 Announce Type: replace-cross Abstract: Mechanistic interpretability often studies Transformer behavior by intervening on internal activations through activation patching, causal tracing, path patching, and steering directions.
arXiv:2608. 02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere.
arXiv:2607. 17696v1 Announce Type: cross Abstract: We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions.
arXiv:2602. 06883v3 Announce Type: replace Abstract: The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversarial robustness.
arXiv:2607. 22367v1 Announce Type: new Abstract: Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers.
arXiv:2502. 00213v5 Announce Type: replace-cross Abstract: Transformers are difficult to optimize with stochastic gradient descent (SGD) and largely rely on adaptive optimizers such as Adam.
arXiv:2607. 25663v1 Announce Type: new Abstract: Transformer adaptation is typically distributed across model depth, even when the intended change is narrow.
arXiv:2605. 27458v2 Announce Type: replace-cross Abstract: Transformer has significantly propelled the development of artificial intelligence, and certainly the development of agents as well.
arXiv:2606. 16694v1 Announce Type: cross Abstract: Transformers are widely used as a general-purpose substrate for learning complex correlations between a large collection of coupled variables, but their internal mechanisms have remained mysterious.
arXiv:2607. 03763v1 Announce Type: cross Abstract: Federated Transformer training increasingly relies on local AdamW, whose adaptive updates can provide much stronger local progress than SGD-based training.
arXiv:2606. 09899v1 Announce Type: cross Abstract: A central goal of mechanistic interpretability is to identify which internal components causally drive a language model's behavior.
arXiv:2607. 18804v1 Announce Type: new Abstract: In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to generate continuations.
arXiv:2603. 19742v2 Announce Type: replace Abstract: Understanding the internal mechanisms of transformer-based large language models (LLMs) is crucial for their reliable deployment and effective operation.