arXiv Machine Learning By Difan Deng, Andreas Bentzen Winje, Lukas Fehring, Marius Lindauer

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models

Read the original on arXiv Machine Learning →

arXiv:2602. 03681v2 Announce Type: replace-cross Abstract: The quadratic computational complexity of softmax transformers has become a bottleneck in long-context scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
3d ago

Switching Linear Attention

Switching Linear Attention (SwiLA) is a new sequence layer that improves upon standard softmax attention by maintaining a fixed-size recurrent state while enhancing representational capacity. It derives its recurrence from a test-time regression framework, using online expectation-maximization in a mixture of linear regressions model. In various benchmarks—including associative recall, in-context language learning, and language modeling—SwiLA achieves strong performance, narrowing the gap to softmax attention and even surpassing it in some settings.

By Hyun Dong Lee, Xavier Gonzalez, Nicolas Zucchet, E. Kelly Buchanan, Emily B. Fox, Scott W. Linderman