arXiv Machine Learning

Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention

arXiv:2608. 11427v1 Announce Type: new Abstract: Full attention exposes every token pair, whereas kernel attention compresses a sequence into a fixed-dimensional sketch.