arXiv Computation and Language

S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

S$^4$R is a method for compressing the Key-Value cache in large language models by building low‑rank subspaces from selectively sampled tokens and performing attention over a sparsely reconstructed KV representation. It initializes key/value bases using a representative prompt subset, reducing reliance on external calibration data while avoiding the high compute cost of full‑prompt decomposition. Experiments on LongBench and RULER with Llama and Qwen models demonstrate up to 5× KV compression with near‑full‑cache accuracy, blending the efficiency of fixed compression with the adaptability of prompt‑dependent approaches.

arXiv AI
3d ago

KV$^2$: A Self-Refining KV Cache

KV$^2$ is a query‑agnostic key‑value cache compression technique that selectively reconstructs only informative in‑context tokens using a lightweight proxy scorer before final eviction scoring. On benchmarks such as RULER, Needle‑in‑a‑Haystack, and LongBench, KV$^2$ outperforms baseline methods, especially under tight memory budgets, achieving higher scores with lower runtime and peak memory than full‑context reconstruction. The approach demonstrates that reusable KV‑cache compression can avoid reprocessing the entire prompt while maintaining quality.

By Johannes Wesch, Danni Liu, Jan Niehues
Hugging Face Trending Papers
Jun 8

End-to-End Context Compression at Scale

Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degrade model quality substantially or require considerable time and compute to compress a single long prompt.

arXiv AI
Jun 9

End-to-End Context Compression at Scale

arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.

By Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov
arXiv Machine Learning
Sep 23

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

CompKV introduces a compensation‑aware sparse attention framework for long‑context LLM inference. It partitions tokens into blocks and optimizes token selection to minimize the error introduced by block‑level mean compensation, using compact block‑level statistics. Experiments on RULER and LongBench‑Pro show CompKV outperforms other sparse baselines and achieves up to a 6.85× speedup over full attention.

By Zhen Huang, Ruizhe Yao, Danyi Liu, Xinrui Chen, Shuwei Li, Siru Zhong, Zijian Cao, Yushan Lai, Mingming Guo, Weijie Zheng, Haohuan Fu
arXiv AI
Jul 8

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

arXiv:2607. 06519v1 Announce Type: new Abstract: Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the layer-specific evidence needed for retrieval and multi-step reasoning.

By Anna C\'ordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jes\'us Olivera