arXiv:2602.04852v3 Announce Type: replace
Abstract: Linear attention offers a computationally efficient yet expressive alternative to softmax attention. However, recent empirical results indicate tha...
By Philipp Nazari, T. Konstantin Rusch
arXiv:2609.13141v1 Announce Type: new
Abstract: Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by selecting a small set of context u...
By Zhiwei Li, Lei Zhu, Hao Gu, Xiang Hu, Yan Wang, Haitao Mi, Sirui Han, Leo Liang, Zhijiang Guo
Switching Linear Attention (SwiLA) is a new sequence layer that improves upon standard softmax attention by maintaining a fixed-size recurrent state while enhancing representational capacity. It derives its recurrence from a test-time regression framework, using online expectation-maximization in a mixture of linear regressions model. In various benchmarks—including associative recall, in-context language learning, and language modeling—SwiLA achieves strong performance, narrowing the gap to softmax attention and even surpassing it in some settings.
By Hyun Dong Lee, Xavier Gonzalez, Nicolas Zucchet, E. Kelly Buchanan, Emily B. Fox, Scott W. Linderman
arXiv:2606. 18587v1 Announce Type: cross Abstract: Decoder-only Transformers compute attention over the KV cache of preceding tokens.
By Zhiyuan Wang, Xuan Luo, Sirui Zeng, Xifeng Yan
arXiv:2606. 02328v1 Announce Type: new Abstract: We explore Riemannian optimization techniques for rank-factored matrix parameters, targeting contemporary deep learning applications.
By Nicholas Knight
arXiv:2609.08615v1 Announce Type: new
Abstract: Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single sca...
By Lenard Dome