arXiv:2606. 01563v1 Announce Type: new Abstract: Autoregressive decoding in Transformer-based language models relies on the KV cache, whose memory footprint grows linearly with sequence length and becomes the primary bottleneck for long-context inference.
By Yu Li, Binxu Li, Tian Lan
arXiv:2607. 27539v1 Announce Type: new Abstract: Exact deletion from persistent language-model memory depends on how that memory represents a record.
By Vishwajith Ramesh
ValueDiff introduces a value‑geometric KV cache eviction strategy for large language models that suppress attention sinks. It ranks tokens by the L2 deviation of their value vectors from the cache mean, a score that aligns with minimal‑disturbance eviction under a max‑entropy assumption. Across several benchmarks—RULER, LongBench, and MATH‑500—ValueDiff consistently retains a higher proportion of useful tokens than prior methods, especially under tight cache budgets.
By Junyoung Park, Jungwook Choi, Mingu Lee
arXiv:2607. 07144v1 Announce Type: new Abstract: The key-value (KV) cache dominates the memory cost of long-context autoregressive inference, and a growing body of work compresses it through quantization, eviction, or offloading.
By Vladimir Gusev
arXiv:2608. 26515v1 Announce Type: cross Abstract: We study online prediction for a specific finite-alphabet, exogenously driven source with infinite input memory.
By Vaneet Aggarwal
The paper introduces Window-Diffusion, a method that accelerates diffusion language model inference by pruning and caching tokens within a sliding window. It categorizes undecoded tokens into active, buffer, and far-field groups, computing only the first two while discarding the rest. Experiments on LLaDA and Dream demonstrate up to 99× speedup with minimal loss in generation quality.
By Fengrui Zuo, Zhiwei Ke, Yiming Liu, Wenqi Lou, Chao Wang, Xuehai Zhou
arXiv:2606. 17872v1 Announce Type: cross Abstract: Large language models (LLMs) outperform earlier architectures on generative inference and long-context tasks, but their large size introduces significant challenges in memory usage, energy cost, and on-device deployment.
By Ning Ni, Yingjie Lao
arXiv:2609.37887v1 Announce Type: new
Abstract: Activation and key-value cache precision change what a quantized language model computes without altering its stored weights. Direct weight-code bounds...
By Arian Eamaz, Mojtaba Soltanalian
arXiv:2604. 25975v2 Announce Type: replace-cross Abstract: Key-Value (KV) caching is essential for large language model inference, yet its memory overhead poses a critical bottleneck for long-context generation.
By Jiaming Yang, Chenwei Tang, Liangli Zhen, Jiancheng Lv
arXiv:2607. 10441v1 Announce Type: cross Abstract: Context engineering decides what information a model carries forward, and current designs meter it in tokens: compressing the past into a bounded recurrent state, keeping a key-value entry for every token, or imposing a fixed budget through a window or eviction rule.
By Siddharth Pal, Viktoria Rojkova
PAGE is a partition‑aware gated KV‑cache eviction method that reframes eviction as a per‑input admission decision. It uses a single label‑free scalar— the early‑to‑late drop in pairwise top‑k head agreement—to classify inputs into a capacity‑bound class (where eviction is catastrophic) and a dilution‑prone class (where eviction is safe or beneficial). By thresholding this drop, PAGE applies a base evictor only when necessary, reducing the harm rate in the capacity‑bound regime from 0.75 to 0.026 and achieving a 29× improvement across four models and benchmarks without retraining the evictor.
By Pankaj Kumar, Subhankar Mishra
The paper investigates how block‑diffusion language models can use a constant‑size cache to enable efficient parallel decoding. By employing sequence mixers that summarize completed blocks into a reusable state and a block‑causal training objective, the authors pretrain three 3B block‑diffusion denoisers (attention, Mamba, and hybrid) on 300 B tokens. The resulting state‑space cache remains O(1) in memory and latency regardless of context length, yielding significant speed‑up and memory savings compared to traditional attention‑based caches, especially at very long sequences.
By Vaibhav Singh, Pierre-Andr\'e No\"el, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko