PolyKV: Heterogeneous Retention and Allocation for KV Cache Compression
arXiv:2606. 15157v1 Announce Type: cross Abstract: KV cache compression is essential for reducing the memory cost of long-context large language model inference.
arXiv:2607. 01520v1 Announce Type: new Abstract: Transformer inference on long sequences is expensive because softmax attention repeatedly reads from a large KV cache.
arXiv:2606. 15157v1 Announce Type: cross Abstract: KV cache compression is essential for reducing the memory cost of long-context large language model inference.
arXiv:2502. 16886v4 Announce Type: replace-cross Abstract: To reduce memory consumption during LLM inference, a handful of methods have been proposed for KV cache pruning.
arXiv:2607. 06519v1 Announce Type: new Abstract: Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the layer-specific evidence needed for retrieval and multi-step reasoning.
arXiv:2608.23843v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compressio...
arXiv:2609.37988v1 Announce Type: new Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This i...
The paper investigates how chunked KV‑cache compression, which reduces memory and attention costs by grouping consecutive tokens, introduces a new positional coordinate called a token’s phase. It finds that models using this compression exhibit phase sensitivity: retrieval performance can vary by up to 40 percentage points depending on a token’s phase, creating periodic weak spots that average benchmark scores hide. By pretraining transformers with different compression designs and applying causal interventions, the authors show that attention components specialize asymmetrically for different phases, and theoretical analysis suggests gradient dynamics may favor such specialization.
arXiv:2608. 07001v1 Announce Type: new Abstract: As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck.
arXiv:2608. 14191v1 Announce Type: new Abstract: The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference.
arXiv:2608. 08684v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging.
EchoPress is a training‑free method for pruning key‑value caches in large language models. It approximates the reconstruction attention used by KVzip by leveraging queries and keys from standard prefill, reconstructing only the first chunk to calibrate importance scores for the rest of the context. Experiments on LongBench and RULER with Qwen3‑8B and Llama‑3.1‑8B‑Instruct show that EchoPress matches KVzip’s task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by 1.7–19.6× and total prefill time by up to 2.9×.
arXiv:2607. 27600v1 Announce Type: new Abstract: Key-value (KV) cache management through compression and eviction strategies has emerged as an important research direction in recent years.
arXiv:2606. 01563v1 Announce Type: new Abstract: Autoregressive decoding in Transformer-based language models relies on the KV cache, whose memory footprint grows linearly with sequence length and becomes the primary bottleneck for long-context inference.