WildCat: Near-Linear Attention in Theory and Practice
arXiv:2602. 10056v2 Announce Type: replace Abstract: We introduce WildCat, a high-accuracy, low-cost approach to compressing the attention mechanism in neural networks.
arXiv:2602. 10056v2 Announce Type: replace Abstract: We introduce WildCat, a high-accuracy, low-cost approach to compressing the attention mechanism in neural networks.
arXiv:2601. 03043v4 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate strong capabilities across a wide range of complex tasks and are increasingly deployed at scale, placing significant demands on inference efficiency.
The paper proves that deep residual self‑attention networks can universally interpolate between any two collections of sequences using only two fixed single‑head attention blocks with Gaussian‑initialized projections. The interpolation is achieved by varying the order, signs, and durations of these blocks, independent of the specific input and output sequences. The result holds for both continuous and finite depth, and the authors also extend the analysis to causal‑masked settings.
FlashBoB introduces an I/O‑efficient algorithm for exact backward‑over‑backward (BoB) in softmax attention, enabling precise second‑order differentiation without large intermediate tensors. By exploiting a hierarchical affine structure, the method confines computation to on‑chip tiles and limits off‑chip memory traffic, achieving θ(N² d²/M) HBM usage. Experiments show FlashBoB scales to sequence lengths of 262K on a single A100 GPU, outperforming prior exact baselines and FlashBack by up to 6.3×.
arXiv:2606. 01294v1 Announce Type: cross Abstract: Linear attention reduces the quadratic cost of softmax attention by maintaining a recurrent fast-weight state, but it consistently lags on in-context retrieval and long-context tasks.
arXiv:2609. 31093v1 Announce Type: new Abstract: Scaling language models to long contexts is limited by the quadratic cost of self-attention.
arXiv:2601. 12145v3 Announce Type: replace Abstract: Softmax attention struggles with long contexts due to structural limitations: the strict sum-to-one constraint forces attention sinks on irrelevant tokens, and probability mass disperses as sequence lengths increase.
arXiv:2605. 16928v2 Announce Type: replace-cross Abstract: Long-context inference in large language models is bottlenecked by the quadratic cost of full attention.
arXiv:2602.04852v3 Announce Type: replace Abstract: Linear attention offers a computationally efficient yet expressive alternative to softmax attention. However, recent empirical results indicate tha...
arXiv:2607. 18759v1 Announce Type: new Abstract: Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not.
arXiv:2602. 03681v2 Announce Type: replace-cross Abstract: The quadratic computational complexity of softmax transformers has become a bottleneck in long-context scenarios.
arXiv:2604. 24432v2 Announce Type: replace-cross Abstract: Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agentic intelligence and recommendation system.