Feed-Forward Steering in Transformer Residual Dynamics
arXiv:2608. 02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere.
arXiv:2605. 11007v2 Announce Type: replace-cross Abstract: We show that the core components of the Transformer -- attention, residual connections, and normalization -- arise naturally from a single geometric state estimation problem.
arXiv:2608. 02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere.
arXiv:2609.15975v1 Announce Type: cross Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...
The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.
arXiv:2609.37717v1 Announce Type: new Abstract: Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state thro...
arXiv:2606. 30440v1 Announce Type: cross Abstract: We present a complete formal proof that transformer architectures, when their internal update mechanisms satisfy a Bayes joint-distribution condition, implement exact Bayesian posterior inference.
arXiv:2607. 15819v1 Announce Type: cross Abstract: In-context learning is a remarkable property of transformers and has recently received a lot of interest.
arXiv:2605. 08475v3 Announce Type: replace-cross Abstract: In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass.
arXiv:2605. 27458v2 Announce Type: replace-cross Abstract: Transformer has significantly propelled the development of artificial intelligence, and certainly the development of agents as well.
arXiv:2601. 22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures.
arXiv:2602. 06883v3 Announce Type: replace Abstract: The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversarial robustness.
arXiv:2606. 17830v1 Announce Type: cross Abstract: Neural network parameter spaces are inherently non-injective, as distinct parameter configurations can realize identical functions through functional equivalence.
arXiv:2606. 16694v1 Announce Type: cross Abstract: Transformers are widely used as a general-purpose substrate for learning complex correlations between a large collection of coupled variables, but their internal mechanisms have remained mysterious.