arXiv Machine Learning

The risk of KV cache compression

arXiv:2607. 01520v1 Announce Type: new Abstract: Transformer inference on long sequences is expensive because softmax attention repeatedly reads from a large KV cache.

arXiv AI
Jul 8

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

arXiv:2607. 06519v1 Announce Type: new Abstract: Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the layer-specific evidence needed for retrieval and multi-step reasoning.

By Anna C\'ordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jes\'us Olivera
arXiv AI
4d ago

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

The paper investigates how chunked KV‑cache compression, which reduces memory and attention costs by grouping consecutive tokens, introduces a new positional coordinate called a token’s phase. It finds that models using this compression exhibit phase sensitivity: retrieval performance can vary by up to 40 percentage points depending on a token’s phase, creating periodic weak spots that average benchmark scores hide. By pretraining transformers with different compression designs and applying causal interventions, the authors show that attention components specialize asymmetrically for different phases, and theoretical analysis suggests gradient dynamics may favor such specialization.

By Xingyu Zhu (Luke), Pu (Luke), Yi, Ziheng Cheng, Ang Lv, Jing Liu, Lexing Ying, Yiyuan Ma, Xin Dong
arXiv Machine Learning
1d ago

EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction

EchoPress is a training‑free method for pruning key‑value caches in large language models. It approximates the reconstruction attention used by KVzip by leveraging queries and keys from standard prefill, reconstructing only the first chunk to calibrate importance scores for the rest of the context. Experiments on LongBench and RULER with Qwen3‑8B and Llama‑3.1‑8B‑Instruct show that EchoPress matches KVzip’s task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by 1.7–19.6× and total prefill time by up to 2.9×.

By Jiawei Lin, Saibo Geng, Thomas Bourgeat