arXiv AI
Jul 8

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

arXiv:2607. 06519v1 Announce Type: new Abstract: Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the layer-specific evidence needed for retrieval and multi-step reasoning.

By Anna C\'ordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jes\'us Olivera
arXiv AI
3d ago

iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD

The paper introduces iS-KV, an online low‑rank KV‑cache compression technique that uses block‑incremental SVD to manage memory during long‑horizon autoregressive decoding. Unlike token‑eviction methods, iS‑KV retains all positions in a compact representation by keeping a recent window exact and incrementally folding older states into bounded‑rank bases, synchronizing coordinates as the basis evolves. Experiments on DeepSeek‑R1‑Distill‑Llama‑8B and Qwen3‑8B show that iS‑KV achieves high accuracy (82.6% and 89.2% respectively) while providing 4.06‑fold and 5.64‑fold compression, outperforming token‑eviction baselines under matched memory budgets.

By Yiren Zhao, Guanghui Song, Tianrui Qin, Kejiang Ye, Cheng-zhong Xu, Xitong Gao
arXiv Machine Learning
Sep 25

A JoLT for the KV cache: Near-Lossless KV Cache Compression via Joint Rank-bit Allocation

The paper introduces JoLT, a training‑free compressor that jointly allocates rank and precision for key‑value (KV) cache compression in long‑context language models. JoLT treats grouped prefill caches as fourth‑order tensors, applies partial Tucker decomposition along token and feature modes, and uses a rotated low‑bit quantizer for residuals, all governed by a single Lagrangian dual under a global byte constraint. Across five models from four architecture families, JoLT achieves 2–3× compression with less than 0.2% perplexity loss, and near‑lossless retrieval accuracy on LLaMA‑3.1‑8B at 64K context up to 3× compression.

By Rahul Krishnan, Volker Schulz