arXiv:2604. 11613v4 Announce Type: replace-cross Abstract: Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque.
By Patrick Lutz, Themistoklis Haris, Arjun Chandra, Aditya Gangrade, Venkatesh Saligrama
arXiv:2601. 17257v2 Announce Type: replace Abstract: We introduce a constrained optimization framework for training transformers that behave like optimization descent algorithms.
By Javier Porras-Valenzuela, Samar Hadou, Alejandro Ribeiro
arXiv:2609.37631v1 Announce Type: new
Abstract: Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work sh...
By Zachary Shinnick, Christian Intern\`o, Hemanth Saratchandran, Anton van den Hengel, Damien Teney
The paper demonstrates that a single-head, single-layer transformer cannot determine whether a bit sequence is ordered, whereas a two-head, single-layer transformer can. This distinction is shown under a model where transformers include an output MLP. The study provides a concrete example of how increasing the number of heads can enhance a transformer’s computational capability.
By Jasper van Doornmalen, Alexander Kozachinskiy, Corinna Mathwieser, Tomasz Steifer, Felipe Urrutia, Jos\'{e} Verschae, Przemys{\l}aw Andrzej Wa{\l}\c{e}ga
arXiv:2509. 12760v5 Announce Type: replace Abstract: We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax activation function, adding Similarity (i.
By Allen Schmaltz
The paper introduces Transformer-Within-Transformer (TWT), a post‑hoc technique that merges contiguous redundant layers in Vision Transformers into a single surrogate layer. By doing so, TWT cuts both parameter count and inference compute while maintaining competitive performance on natural image tasks with only half the depth. In histopathology applications, TWT not only matches but sometimes surpasses the baseline model’s performance.
By Dhananjay Tomar, Marius Aasan, Andreas Kleppe, Ad\'in Ram\'irez Rivera
arXiv:2608. 20134v1 Announce Type: cross Abstract: We present a novel view on feature evolution in Vision Transformers (ViTs) by visualizing the training process over two dimensions -- network depth (layer) and training time (epochs).
By Joonas J\"arve, Halil Ibrahim Aysel, Tarun Khajuria, Meelis Kull
arXiv:2510. 09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.
By Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Krishna Kumar, Markos A. Katsoulakis
arXiv:2410. 11500v2 Announce Type: replace-cross Abstract: In this paper, we establish a collection of covering number bounds for linear function classes under various norm constraints on the inputs and matrices.
By Lan V. Truong
The paper introduces a technique for examining Vision Transformers by decomposing each affine layer’s weight matrix with Singular Value Decomposition and projecting activations onto the leading right singular vectors, yielding compact, layer‑intrinsic representations. By fitting class‑conditional density models at each layer, the authors generate per‑class typicality scores that are stacked into two‑dimensional typicality maps, summarizing how class‑specific evidence evolves through the network. From these maps, two post‑hoc out‑of‑distribution detection scores are derived: the Prototype Alignment Score (PAS), which measures agreement with class reference prototypes, and the Multi‑Layer Soft Voting (MLSV) score, which captures cross‑layer consensus without stored prototypes, achieving competitive performance on ViT‑B/16 fine‑tuned on CIFAR‑100 without retraining or OOD exposure.
By Aldo Sean Sartor, Leandro de Souza Rosa, Andriy Enttsel, Mauro Mangia, Riccardo Rovatti
arXiv:2606. 14555v1 Announce Type: cross Abstract: Modern image classifiers widely adopt global average pooling (GAP) followed by a linear classification head.
By Aray Karjauv