arXiv AI

AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD

arXiv AI
3d ago

iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD

The paper introduces iS-KV, an online low‑rank KV‑cache compression technique that uses block‑incremental SVD to manage memory during long‑horizon autoregressive decoding. Unlike token‑eviction methods, iS‑KV retains all positions in a compact representation by keeping a recent window exact and incrementally folding older states into bounded‑rank bases, synchronizing coordinates as the basis evolves. Experiments on DeepSeek‑R1‑Distill‑Llama‑8B and Qwen3‑8B show that iS‑KV achieves high accuracy (82.6% and 89.2% respectively) while providing 4.06‑fold and 5.64‑fold compression, outperforming token‑eviction baselines under matched memory budgets.

By Yiren Zhao, Guanghui Song, Tianrui Qin, Kejiang Ye, Cheng-zhong Xu, Xitong Gao
arXiv Machine Learning
Sep 25

A JoLT for the KV cache: Near-Lossless KV Cache Compression via Joint Rank-bit Allocation

The paper introduces JoLT, a training‑free compressor that jointly allocates rank and precision for key‑value (KV) cache compression in long‑context language models. JoLT treats grouped prefill caches as fourth‑order tensors, applies partial Tucker decomposition along token and feature modes, and uses a rotated low‑bit quantizer for residuals, all governed by a single Lagrangian dual under a global byte constraint. Across five models from four architecture families, JoLT achieves 2–3× compression with less than 0.2% perplexity loss, and near‑lossless retrieval accuracy on LLaMA‑3.1‑8B at 64K context up to 3× compression.

By Rahul Krishnan, Volker Schulz