The Head Complexity of Boolean Functions in Single-Layer Attention
arXiv:2609. 04046v1 Announce Type: cross Abstract: What can a single layer of self-attention compute?
arXiv:2608. 04243v1 Announce Type: new Abstract: Multi-head attention layers produce vector representations that support multiple downstream tasks.
arXiv:2609. 04046v1 Announce Type: cross Abstract: What can a single layer of self-attention compute?
arXiv:2608. 11427v1 Announce Type: new Abstract: Full attention exposes every token pair, whereas kernel attention compresses a sequence into a fixed-dimensional sketch.
arXiv:2502. 01015v5 Announce Type: replace Abstract: Task arithmetic, representing downstream tasks through linear operations on task vectors, has emerged as a simple yet powerful paradigm for transferring knowledge across diverse settings.
Full attention exposes every token pair, whereas kernel attention compresses a sequence into a fixed-dimensional sketch. We show that this distinction becomes exponential at the first context length containing two competing candidates.
arXiv:2606. 07205v1 Announce Type: cross Abstract: The attention mechanism is a cornerstone of modern transformer architectures.
arXiv:2609.37261v1 Announce Type: new Abstract: Softmax attention is ubiquitous in modern machine learning, but its quadratic scaling with sequence length makes it costly. To reduce this cost, attent...
arXiv:2609.23094v1 Announce Type: cross Abstract: We study the number of prototypes needed to represent Boolean functions by nearest-neighbour classification. There are two distinct settings: the pro...
arXiv:2608. 03294v1 Announce Type: new Abstract: We study the problem of learning multi-head softmax attention from black-box input-output access.
arXiv:2509. 07963v2 Announce Type: replace Abstract: The core component of attention is the scoring function, which transforms the inputs into low-dimensional queries and keys and takes the dot product of each pair.
arXiv:2606. 07713v1 Announce Type: cross Abstract: The attention mechanism is the dominant computational bottleneck in modern transformer-based AI.
arXiv:2609.36760v1 Announce Type: new Abstract: Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memor...
arXiv:2606. 04032v1 Announce Type: cross Abstract: Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role.