arXiv Machine Learning

Specialization of softmax attention heads: insights from the high-dimensional single-location model

arXiv:2603. 03993v2 Announce Type: replace Abstract: Multi-head attention enables transformer models to represent multiple attention patterns simultaneously.

arXiv Machine Learning
Jul 7

Incremental Learning of Sparse Attention Patterns in Transformers

arXiv:2602. 19143v2 Announce Type: replace Abstract: This paper studies simple transformers trained on a high-order Markov chain, where the model must incorporate information from multiple past positions, each with different statistical importance.

By O\u{g}uz Kaan Y\"uksel, Rodrigo Alvarez Lucendo, Nicolas Flammarion