arXiv Machine Learning

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs

arXiv:2607. 12550v1 Announce Type: new Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference.

arXiv Machine Learning
Sep 25

A JoLT for the KV cache: Near-Lossless KV Cache Compression via Joint Rank-bit Allocation

The paper introduces JoLT, a training‑free compressor that jointly allocates rank and precision for key‑value (KV) cache compression in long‑context language models. JoLT treats grouped prefill caches as fourth‑order tensors, applies partial Tucker decomposition along token and feature modes, and uses a rotated low‑bit quantizer for residuals, all governed by a single Lagrangian dual under a global byte constraint. Across five models from four architecture families, JoLT achieves 2–3× compression with less than 0.2% perplexity loss, and near‑lossless retrieval accuracy on LLaMA‑3.1‑8B at 64K context up to 3× compression.

By Rahul Krishnan, Volker Schulz
arXiv AI
Jun 4

Stochastic Sparse Attention for Memory-Bound Inference

arXiv:2605. 01910v2 Announce Type: replace-cross Abstract: Autoregressive decoding becomes bandwidth-limited at long contexts, as generating each token requires reading all $n_k$ key and value vectors from KV cache.

By Kyle Lee, Corentin Delacour, Kevin Callahan-Coray, Kyle Jiang, Can Yaras, Samet Oymak, Tathagata Srimani, Kerem Y. Camsari
arXiv AI
Jun 24

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

arXiv:2606. 24467v1 Announce Type: new Abstract: Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware.

By Xiaolin Lin, Jingcun Wang, Olga Kondrateva, Yiyu Shi, Bing Li, Grace Li Zhang