arXiv:2606. 29563v1 Announce Type: cross Abstract: Large language models (LLMs) excel at complex tasks like question answering and summarization, thanks to their ability to handle long-context inputs.
By Shuvendu Roy, Mengyao Zhai, Hossein Hajimirsadeghi, Golnoosh Samei
arXiv:2602. 10238v2 Announce Type: replace-cross Abstract: The growing size of Large Language Models (LLMs) makes efficient inference challenging, primarily due to the memory demands of the autoregressive Key-Value (KV) cache.
By Luca Moschella, Laura Manduchi, Ozan Sener
The paper introduces EpiKV, an epiphany‑aware key–value cache eviction strategy that avoids using the attention matrix. It leverages hidden‑state shifts and recent query–key relevance to rank cached tokens, matching or surpassing the performance of existing attention‑based eviction methods while remaining compatible with fast inference kernels. Experiments on multiple benchmarks show that EpiKV improves inference throughput without sacrificing accuracy.
By Steven Kolawole, Virginia Smith
arXiv:2604.05012v2 Announce Type: replace-cross
Abstract: Efficient inference with Large Language Models (LLMs) increasingly relies on Key-Value (KV) caches to store previously computed key and value...
By Oteo Mamo, Olga Kogiou, Hyunjin Yi, Weikuan Yu
arXiv:2606. 24467v1 Announce Type: new Abstract: Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware.
By Xiaolin Lin, Jingcun Wang, Olga Kondrateva, Yiyu Shi, Bing Li, Grace Li Zhang
arXiv:2609.37988v1 Announce Type: new
Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This i...
By Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi