arXiv Machine Learning

Faster Query-Key Learning Sharpens Attention in Self-Attention Models

arXiv:2608. 06776v1 Announce Type: new Abstract: A standard self-attention layer consists of two interacting circuits: the query-key circuit that governs attention allocation, and the output-value circuit that maps attended representations to predictions.