Training-Free Universal Approximation by Prompting Random Transformers
arXiv:2608. 09558v1 Announce Type: new Abstract: How expressive is prompting a transformer?
arXiv:2512. 11784v2 Announce Type: replace Abstract: Softmax attention is a central component of transformer architectures, yet its nonlinear structure poses significant challenges for theoretical analysis.
arXiv:2608. 09558v1 Announce Type: new Abstract: How expressive is prompting a transformer?
arXiv:2607. 00479v1 Announce Type: new Abstract: Transformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning.
arXiv:2607. 23050v1 Announce Type: new Abstract: Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the minimum model capacity required to solve it?
arXiv:2607. 18759v1 Announce Type: new Abstract: Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not.
arXiv:2606. 22406v2 Announce Type: replace Abstract: Attention mechanisms have demonstrated remarkable empirical success in identifying relevant information from large collections of tokens, yet the theoretical principles underlying this behavior remain poorly understood.
arXiv:2602. 03681v2 Announce Type: replace-cross Abstract: The quadratic computational complexity of softmax transformers has become a bottleneck in long-context scenarios.
The paper proves that deep residual self‑attention networks can universally interpolate between any two collections of sequences using only two fixed single‑head attention blocks with Gaussian‑initialized projections. The interpolation is achieved by varying the order, signs, and durations of these blocks, independent of the specific input and output sequences. The result holds for both continuous and finite depth, and the authors also extend the analysis to causal‑masked settings.
arXiv:2512. 10656v3 Announce Type: replace Abstract: As context windows in large language models continue to expand, it is essential to characterize how attention behaves at extreme sequence lengths.
arXiv:2605. 08475v3 Announce Type: replace-cross Abstract: In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass.
arXiv:2507. 07814v2 Announce Type: replace Abstract: We introduce a novel upper bound on the local Lipschitz constant of the dot-product self-attention block showing its dependence on the attention map distributions.
arXiv:2601. 12145v3 Announce Type: replace Abstract: Softmax attention struggles with long contexts due to structural limitations: the strict sum-to-one constraint forces attention sinks on irrelevant tokens, and probability mass disperses as sequence lengths increase.
Switching Linear Attention (SwiLA) is a new sequence layer that improves upon standard softmax attention by maintaining a fixed-size recurrent state while enhancing representational capacity. It derives its recurrence from a test-time regression framework, using online expectation-maximization in a mixture of linear regressions model. In various benchmarks—including associative recall, in-context language learning, and language modeling—SwiLA achieves strong performance, narrowing the gap to softmax attention and even surpassing it in some settings.