Many Transformer explainers start with the finished architecture. We ask why it looks the way it does.
By Sankar Srinivasan
arXiv:2606. 17830v1 Announce Type: cross Abstract: Neural network parameter spaces are inherently non-injective, as distinct parameter configurations can realize identical functions through functional equivalence.
By Viet-Hoang Tran, Vinh Khanh Bui, Van-Hoan Trinh, Tan Lai Ngoc, Tan M. Nguyen
arXiv:2606. 06160v1 Announce Type: new Abstract: RoPE-trained transformers distinguish absolute position in their attention patterns, even though RoPE encodes only relative offsets in the inner product.
By Valeria Ruscio, Umberto Nanni, Fabrizio Silvestri
arXiv:2609.37921v1 Announce Type: new
Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...
By Erkan Turan, Gaspard Abel, Maks Ovsjanikov
t0-alpha is a decoder-style patch transformer for probabilistic time-series forecasting. Raw series are split into 32-step patches, embedded, processed through causal time-attention and group-attention layers, and decoded into future quantiles rather than a single point forecast.
By Sean Moran