arXiv Machine Learning

Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention

arXiv:2608. 11427v1 Announce Type: new Abstract: Full attention exposes every token pair, whereas kernel attention compresses a sequence into a fixed-dimensional sketch.

arXiv Computation and Language
3d ago

Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head

The paper introduces NAMOH, a native sparse attention mechanism that activates only a subset of heads per token, allowing each head to attend to a limited subsequence of tokens. By scaling the number of heads while keeping the active heads per token fixed, the method shortens head histories and reduces key‑value access without increasing overall storage. Experiments demonstrate that NAMOH can outperform fully activated models with the same parameter count and enable more efficient long‑context inference than smaller dense models.

By Zizhuo Fu, Runsheng Wang, Meng Li