The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.
By Ke Wan, Chen Chen
arXiv:2609.36337v1 Announce Type: new
Abstract: Tabular foundation models achieve strong performance by conditioning on labelled examples in context, but softmax attention limits their use on large d...
By David Schnurr, Felix Sarnthein, Thomas Hofmann, Imanol Schlag
arXiv:2604.19021v3 Announce Type: replace
Abstract: Linear attention mechanisms have emerged as promising alternatives to softmax attention, offering linear-time complexity during inference. Recent a...
By Pingwei Sun, Yuxuan Hu, Jianchao Tan, Xue Wang, Jiaqi Zhang, Yifan Lu, Yerui Sun, Yuchen Xie, Xunliang Cai
arXiv:2606. 27791v1 Announce Type: cross Abstract: Hybrid attention models that mix full and sliding-window attention across layers offer a promising approach to efficient long-context inference, but the critical question of \emph{which layers} should retain full attention remains unsolved.
By Qiong Tang, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao
arXiv:2606. 01294v1 Announce Type: cross Abstract: Linear attention reduces the quadratic cost of softmax attention by maintaining a recurrent fast-weight state, but it consistently lags on in-context retrieval and long-context tasks.
By Dong Le, Thong Nguyen, Cong-Duy Nguyen, Anh Tuan Luu
The paper introduces Recency Forcing, a technique that addresses the long‑horizon degradation in autoregressive video generation caused by KV eviction mismatch. By applying a timestep‑dependent bias—Temporal Response Bias—derived from a positional response measure, the method reduces the influence of distant frames during inference without altering context length or training objectives. An exact reformulation, Biased Attention Reparameterization, enables this bias to be applied as a standard FlashAttention call with zero overhead, achieving state‑of‑the‑art long‑horizon generation quality on VBench datasets.
By Tri Cao, Hung Nguyen, Phong Nguyen, Khoi Nguyen
The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While numerous post hoc linearization pipelines exist, it is difficult to identify which components preserve model quality.
The paper introduces NAMOH, a native sparse attention mechanism that activates only a subset of heads per token, allowing each head to attend to a limited subsequence of tokens. By scaling the number of heads while keeping the active heads per token fixed, the method shortens head histories and reduces key‑value access without increasing overall storage. Experiments demonstrate that NAMOH can outperform fully activated models with the same parameter count and enable more efficient long‑context inference than smaller dense models.
By Zizhuo Fu, Runsheng Wang, Meng Li
arXiv:2608. 01672v1 Announce Type: cross Abstract: Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later.
By Zixuan Wang, Xingyu Dang, Rui-Jie Zhu, Zixin Wen, Hengyu Fu, Wenhao Chai, Jason D. Lee
arXiv:2505. 15548v2 Announce Type: replace Abstract: Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic.
By Suvadeep Hajra
arXiv:2607. 05583v1 Announce Type: cross Abstract: Contemporary language models are dominated by the transformer architecture, which leverages self-attention mechanisms to enable more efficient, parallelized training across a wide set of documents and corpora.
By Archie Chaudhury
arXiv:2509.12635v4 Announce Type: replace-cross
Abstract: We prove under practical assumptions that Rotary Positional Embedding (RoPE) introduces an intrinsic distance-dependent bias in attention sco...
By Yu Wang, Sheng Shen, R\'emi Munos, Hongyuan Zhan, Yuandong Tian