arXiv:2608.23843v1 Announce Type: new
Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compressio...
By Zizhong Wang, Jieying Wang, Zhao Zhang, Jiajia Li
arXiv:2610.02953v1 Announce Type: new
Abstract: Long-context LLM serving is increasingly bottlenecked by KV-cache memory, especially in resource-constrained scenarios. Among existing KV-cache compres...
By Zihan Teng, Jiayu Zhao, Wentao Ren, Minhao Fan, Tianrui Ma, Song Chen, Weichen Liu
arXiv:2607. 16213v1 Announce Type: new Abstract: Large Language Models (LLMs) generate text autoregressively, relying on a key-value (KV) cache whose memory footprint grows linearly with context length, creating a major bottleneck.
By Soumia Bouyahiaoui, Manel Kara laouar, Aicha Boutorh, Mohamed Hadj Ameur
arXiv:2604.11501v2 Announce Type: replace-cross
Abstract: Rank reduction discards dimensions; quantization keeps them at lower precision. Comparing the two requires a choice of what compression shoul...
By Samuel Salfati
arXiv:2607. 24331v1 Announce Type: new Abstract: As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually becomes a significant bottleneck as the context window continues to grow.
By Tan T. Nguyen, Quan V. Dang
arXiv:2602. 10238v2 Announce Type: replace-cross Abstract: The growing size of Large Language Models (LLMs) makes efficient inference challenging, primarily due to the memory demands of the autoregressive Key-Value (KV) cache.
By Luca Moschella, Laura Manduchi, Ozan Sener
arXiv:2607. 15498v1 Announce Type: cross Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference.
By Shahrzad Esmat, Dhawal Shah, Ali Jannesari
The paper introduces iS-KV, an online low‑rank KV‑cache compression technique that uses block‑incremental SVD to manage memory during long‑horizon autoregressive decoding. Unlike token‑eviction methods, iS‑KV retains all positions in a compact representation by keeping a recent window exact and incrementally folding older states into bounded‑rank bases, synchronizing coordinates as the basis evolves. Experiments on DeepSeek‑R1‑Distill‑Llama‑8B and Qwen3‑8B show that iS‑KV achieves high accuracy (82.6% and 89.2% respectively) while providing 4.06‑fold and 5.64‑fold compression, outperforming token‑eviction baselines under matched memory budgets.
By Yiren Zhao, Guanghui Song, Tianrui Qin, Kejiang Ye, Cheng-zhong Xu, Xitong Gao
The paper introduces JoLT, a training‑free compressor that jointly allocates rank and precision for key‑value (KV) cache compression in long‑context language models. JoLT treats grouped prefill caches as fourth‑order tensors, applies partial Tucker decomposition along token and feature modes, and uses a rotated low‑bit quantizer for residuals, all governed by a single Lagrangian dual under a global byte constraint. Across five models from four architecture families, JoLT achieves 2–3× compression with less than 0.2% perplexity loss, and near‑lossless retrieval accuracy on LLaMA‑3.1‑8B at 64K context up to 3× compression.
By Rahul Krishnan, Volker Schulz
arXiv:2607. 05061v1 Announce Type: new Abstract: Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length.
By Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl, David Stap, Thomas Schmied, Sebastian B\"ock, G\"unter Klambauer, Sepp Hochreiter
arXiv:2610.02235v1 Announce Type: cross
Abstract: Long-context decoding is increasingly constrained by key--value (KV) cache memory and bandwidth. Existing fixed-budget compression methods typically...
By Shuxin Liu, Qing Liu, Yi Du, Ou Wu
arXiv:2609.13285v1 Announce Type: cross
Abstract: The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with sequence length. Grouped-query a...
By Vishesh Tripathi, Abhay Kumar, Ramsha Khan