arXiv Machine Learning
Sep 23

Transformer Heads Looking for Order

The paper demonstrates that a single-head, single-layer transformer cannot determine whether a bit sequence is ordered, whereas a two-head, single-layer transformer can. This distinction is shown under a model where transformers include an output MLP. The study provides a concrete example of how increasing the number of heads can enhance a transformer’s computational capability.

By Jasper van Doornmalen, Alexander Kozachinskiy, Corinna Mathwieser, Tomasz Steifer, Felipe Urrutia, Jos\'{e} Verschae, Przemys{\l}aw Andrzej Wa{\l}\c{e}ga
arXiv Machine Learning
Jun 9

Similarity-Distance-Magnitude Activations

arXiv:2509. 12760v5 Announce Type: replace Abstract: We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax activation function, adding Similarity (i.

By Allen Schmaltz