arXiv Machine Learning

DeltaLog: Deferred Materialization of Recurrent States for Linear Attention Decoding

arXiv:2608. 15533v1 Announce Type: cross Abstract: Linear attention models eliminate the quadratic prefix computation and context-growing KV cache of softmax attention by replacing pairwise token interactions with recurrent state updates.

arXiv Machine Learning
Jun 10

SinkRec: Mitigating Semantic State Sink in Long Sequence Recommendation with Memory-Conditioned Gated Delta Networks

arXiv:2606. 09888v1 Announce Type: new Abstract: Linear attention provides an efficient backbone for long-sequence recommendation by avoiding the quadratic cost of standard Transformers, but its compressed recurrent state can be dominated by repetitive behavior patterns.

By Zhuang Zhuang, Zhipeng Wei, Ji Dai, Jie Chen, Fei Pan, Peng Jiang, Kun Gai
arXiv Machine Learning
Sep 24

DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video

DeltaS is a query‑agnostic, training‑free method for evicting key‑value cache entries in hybrid video‑language models that combine linear and full attention. It uses the change in the recurrent state of gated‑delta linear attention—called state drift—to decide which video chunks to keep, selecting those that induce larger normalized state changes. In experiments with a fixed memory budget, DeltaS outperforms position‑, attention‑, and key‑value‑based eviction signals, improving performance by 2.1 points on average across six long‑video benchmarks and 5.6 points on the longest benchmark, while adding only 1.9% of the forward‑pass cost.

By Taeyoun Kwon, Seungjin Kim, Hyeonyu Kim, Moon Hwan Kim
arXiv Machine Learning
Sep 21

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

Elastic Threshold Attention (ETA) is a trainable attention mechanism that dynamically predicts contextual thresholds from query representations, enabling selective pruning of KV cache tokens during long‑context decoding. By multiplicatively suppressing sub‑threshold logits during training, ETA avoids representation collapse and eliminates localized attention sinks, allowing a 1.45B model to match dense attention performance at roughly 85% training sparsity and 38% active decode density. At inference, a custom Triton kernel achieves up to 2.5× faster decoding on sequences up to 512K tokens, and an offline calibration step can further reduce compute by 27% by freezing per‑head thresholds.

By Themistoklis Haris, Henry Li, Maryam Karimzadehgan
arXiv Machine Learning
Aug 4

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

arXiv:2608. 01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound.

By Li Wang, Yi Su, Xiabao Wu, Chiran You, Yongchao Liu, Zhan Qiu, Juelu Zhang, Jiajun Zheng, Fangxin Liu, Jie Zhang, Chen Tian, Chengying Huan
arXiv Machine Learning
Sep 24

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.

By Ke Wan, Chen Chen