arXiv AI

Q-Delta: Beyond Key-Value Associative State Evolution

arXiv:2606. 08804v1 Announce Type: new Abstract: Linear attention reformulates sequence modeling as recurrent state evolution, enabling efficient linear-time inference.

arXiv Machine Learning
Jun 10

SinkRec: Mitigating Semantic State Sink in Long Sequence Recommendation with Memory-Conditioned Gated Delta Networks

arXiv:2606. 09888v1 Announce Type: new Abstract: Linear attention provides an efficient backbone for long-sequence recommendation by avoiding the quadratic cost of standard Transformers, but its compressed recurrent state can be dominated by repetitive behavior patterns.

By Zhuang Zhuang, Zhipeng Wei, Ji Dai, Jie Chen, Fei Pan, Peng Jiang, Kun Gai
arXiv Machine Learning
Aug 4

Structured Recurrent Mixers for Massively Parallelized Sequence Generation

arXiv:2605. 08696v4 Announce Type: replace-cross Abstract: Over the last two decades, language modeling has experienced a shift from the use of predominantly recurrent architectures that process tokens sequentially during training and inference to non-recurrent models that process sequence elements in parallel during training, which results in greater training efficiency and stability at the expense of lower inference throughput.

By Benjamin L. Badger
arXiv Machine Learning
Sep 24

DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video

DeltaS is a query‑agnostic, training‑free method for evicting key‑value cache entries in hybrid video‑language models that combine linear and full attention. It uses the change in the recurrent state of gated‑delta linear attention—called state drift—to decide which video chunks to keep, selecting those that induce larger normalized state changes. In experiments with a fixed memory budget, DeltaS outperforms position‑, attention‑, and key‑value‑based eviction signals, improving performance by 2.1 points on average across six long‑video benchmarks and 5.6 points on the longest benchmark, while adding only 1.9% of the forward‑pass cost.

By Taeyoun Kwon, Seungjin Kim, Hyeonyu Kim, Moon Hwan Kim
arXiv Computation and Language
Sep 4

SGD-KV: Summarization Guided KV Cache Compression

SGD-KV is a head‑aware framework for compressing key‑value caches in large language models. It uses a chunk‑summarization diagnostic task to identify attention heads that specialize in hierarchical information aggregation, allowing the KV cache budget to be allocated based on each head’s summarization score. Experiments on Qwen2.5‑7B‑1M and Qwen3‑32B show state‑of‑the‑art performance on up to 1M‑token contexts while cutting KV cache memory usage by up to 75%.

By Zeyu Liu, Woomin Song, Xuandi Fu, Sai Muralidhar Jayanthi, Vivek Govindan, Aram Galstyan, Sravan Babu Bodapati, Srikanth Ronanki
Hugging Face Trending Papers
Sep 3

SGD-KV: Summarization Guided KV Cache Compression

SGD-KV is a head‑aware framework that compresses key‑value caches in large language models by using a chunk‑summarization diagnostic task to identify attention heads that specialize in hierarchical information aggregation. It prioritizes these heads during compression, achieving state‑of‑the‑art performance on long‑context benchmarks with up to 1M tokens while cutting KV cache memory usage by as much as 75%. Experiments on Qwen2.5‑7B‑1M and Qwen3‑32B confirm that allocating cache budget based on summarization scores yields a superior efficiency‑accuracy trade‑off for long‑context inference.