arXiv Machine Learning

One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context

arXiv Machine Learning
Sep 23

Transformer Heads Looking for Order

The paper demonstrates that a single-head, single-layer transformer cannot determine whether a bit sequence is ordered, whereas a two-head, single-layer transformer can. This distinction is shown under a model where transformers include an output MLP. The study provides a concrete example of how increasing the number of heads can enhance a transformer’s computational capability.

By Jasper van Doornmalen, Alexander Kozachinskiy, Corinna Mathwieser, Tomasz Steifer, Felipe Urrutia, Jos\'{e} Verschae, Przemys{\l}aw Andrzej Wa{\l}\c{e}ga
arXiv Machine Learning
Jun 9

Similarity-Distance-Magnitude Activations

arXiv:2509. 12760v5 Announce Type: replace Abstract: We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax activation function, adding Similarity (i.

By Allen Schmaltz
arXiv Computer Vision
Sep 18

A Smaller Transformer in Your Transformer

The paper introduces Transformer-Within-Transformer (TWT), a post‑hoc technique that merges contiguous redundant layers in Vision Transformers into a single surrogate layer. By doing so, TWT cuts both parameter count and inference compute while maintaining competitive performance on natural image tasks with only half the depth. In histopathology applications, TWT not only matches but sometimes surpasses the baseline model’s performance.

By Dhananjay Tomar, Marius Aasan, Andreas Kleppe, Ad\'in Ram\'irez Rivera
arXiv Computer Vision
Aug 25

SVD-Based Typicality Maps for Out-of-Distribution Detection in Vision Transformers

The paper introduces a technique for examining Vision Transformers by decomposing each affine layer’s weight matrix with Singular Value Decomposition and projecting activations onto the leading right singular vectors, yielding compact, layer‑intrinsic representations. By fitting class‑conditional density models at each layer, the authors generate per‑class typicality scores that are stacked into two‑dimensional typicality maps, summarizing how class‑specific evidence evolves through the network. From these maps, two post‑hoc out‑of‑distribution detection scores are derived: the Prototype Alignment Score (PAS), which measures agreement with class reference prototypes, and the Multi‑Layer Soft Voting (MLSV) score, which captures cross‑layer consensus without stored prototypes, achieving competitive performance on ViT‑B/16 fine‑tuned on CIFAR‑100 without retraining or OOD exposure.

By Aldo Sean Sartor, Leandro de Souza Rosa, Andriy Enttsel, Mauro Mangia, Riccardo Rovatti