Hugging Face Trending Papers

Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention

Full attention exposes every token pair, whereas kernel attention compresses a sequence into a fixed-dimensional sketch. We show that this distinction becomes exponential at the first context length containing two competing candidates.