Before Q, K, and V: Reconstructing the Transformer
Many Transformer explainers start with the finished architecture. We ask why it looks the way it does.
The article explains how transformers, which rely on self‑attention, can lose the natural order of time‑series data when fed scalar observations. It discusses the role of positional encoding in re‑introducing sequence order and provides a visual guide to illustrate this concept.
Many Transformer explainers start with the finished architecture. We ask why it looks the way it does.
arXiv:2606. 17830v1 Announce Type: cross Abstract: Neural network parameter spaces are inherently non-injective, as distinct parameter configurations can realize identical functions through functional equivalence.
arXiv:2606. 06160v1 Announce Type: new Abstract: RoPE-trained transformers distinguish absolute position in their attention patterns, even though RoPE encodes only relative offsets in the inner product.
arXiv:2609.37921v1 Announce Type: new Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...
t0-alpha is a decoder-style patch transformer for probabilistic time-series forecasting. Raw series are split into 32-step patches, embedded, processed through causal time-attention and group-attention layers, and decoded into future quantiles rather than a single point forecast.
arXiv:2602.01605v2 Announce Type: replace Abstract: Time Series Foundation Models (TSFMs) leverage extensive pretraining to accurately predict unseen time series during inference, without the need fo...
The paper investigates distance generalization in transformer models, focusing on how well they can handle changes in inter-token distances between training and inference while keeping context length fixed. Using two synthetic delay-copy tasks that require copying tokens after finite delays, the authors evaluate the impact of positional encoding schemes (RoPE, ALiBi, and NoPE), the diversity of distances seen during training, and the conditions under which distance transfer learning is beneficial or detrimental. Their comprehensive study highlights the importance of understanding the underlying mechanisms that govern distance generalization in transformers.
arXiv:2509. 10534v3 Announce Type: replace-cross Abstract: The attention mechanism in a Transformer architecture matches key to query based on both content -- the what -- and position in a sequence -- the where.
arXiv:2512. 21113v2 Announce Type: replace Abstract: Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective.
arXiv:2606. 01532v1 Announce Type: new Abstract: Positional encoding (PE) is widely viewed as necessary for transformers to process ordered sequences: without them, the next-token map appears permutation-invariant in its context tokens.