arXiv Computation and Language By Jialong Han, You Wu, Kewei Tu

S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

Read the original on arXiv Computation and Language →

S$^4$R is a method for compressing the Key-Value cache in large language models by building low‑rank subspaces from selectively sampled tokens and performing attention over a sparsely reconstructed KV representation. It initializes key/value bases using a representative prompt subset, reducing reliance on external calibration data while avoiding the high compute cost of full‑prompt decomposition. Experiments on LongBench and RULER with Llama and Qwen models demonstrate up to 5× KV compression with near‑full‑cache accuracy, blending the efficiency of fixed compression with the adaptability of prompt‑dependent approaches.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
3d ago

KV$^2$: A Self-Refining KV Cache

KV$^2$ is a query‑agnostic key‑value cache compression technique that selectively reconstructs only informative in‑context tokens using a lightweight proxy scorer before final eviction scoring. On benchmarks such as RULER, Needle‑in‑a‑Haystack, and LongBench, KV$^2$ outperforms baseline methods, especially under tight memory budgets, achieving higher scores with lower runtime and peak memory than full‑context reconstruction. The approach demonstrates that reusable KV‑cache compression can avoid reprocessing the entire prompt while maintaining quality.

By Johannes Wesch, Danni Liu, Jan Niehues