One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2604. 11613v4 Announce Type: replace-cross Abstract: Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque.
arXiv:2601. 17257v2 Announce Type: replace Abstract: We introduce a constrained optimization framework for training transformers that behave like optimization descent algorithms.
arXiv:2609.37631v1 Announce Type: new Abstract: Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work sh...
The paper demonstrates that a single-head, single-layer transformer cannot determine whether a bit sequence is ordered, whereas a two-head, single-layer transformer can. This distinction is shown under a model where transformers include an output MLP. The study provides a concrete example of how increasing the number of heads can enhance a transformer’s computational capability.
arXiv:2509. 12760v5 Announce Type: replace Abstract: We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax activation function, adding Similarity (i.