arXiv:2607. 21475v1 Announce Type: cross Abstract: Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest.
By Peng Xie
arXiv:2609.27981v1 Announce Type: cross
Abstract: KV-cache eviction is typically evaluated through average quality-memory trade-offs, yet a small average loss can hide requests whose utility degrades...
By Beomgu Kang, SoJin Yun, Hojoon Kim, Hyunseok Seo
arXiv:2608.28293v1 Announce Type: new
Abstract: The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible...
By Renato Geh, Alex Chen, Daniel Israel, Aditya Grover, Guy Van den Broeck
The paper introduces Random Attention, a method that evicts KV cache entries uniformly at random within each attention head while preserving the prompt. Experiments on four models and six reasoning tasks show that this simple strategy matches the performance of the best existing eviction methods and achieves 32‑43% higher throughput in vLLM deployments. The authors explain that the prompt is the most fragile cache component and that reasoning traces are redundantly stored across text and attention heads, making a selection score unnecessary.
By Heng Wang, Jielin Qiu, Wenting Zhao, Cheng Qian, Liangwei Yang, Jiawei Han, Heng Ji, Silvio Savarese, Shelby Heinecke, Huan Wang
PAGE is a partition‑aware gated KV‑cache eviction method that reframes eviction as a per‑input admission decision. It uses a single label‑free scalar— the early‑to‑late drop in pairwise top‑k head agreement—to classify inputs into a capacity‑bound class (where eviction is catastrophic) and a dilution‑prone class (where eviction is safe or beneficial). By thresholding this drop, PAGE applies a base evictor only when necessary, reducing the harm rate in the capacity‑bound regime from 0.75 to 0.026 and achieving a 29× improvement across four models and benchmarks without retraining the evictor.
By Pankaj Kumar, Subhankar Mishra
The paper introduces Random Attention, a method that evicts KV cache entries uniformly at random within each attention head while preserving the prompt. It demonstrates that this simple strategy matches or surpasses more complex eviction schemes across four models and six reasoning tasks, achieving 32‑43% higher throughput in vLLM deployments. Experiments reveal that the prompt is the most fragile cache component and that redundancy in the reasoning trace across text and attention heads protects against random eviction, eliminating the need for a selection score.