arXiv:2608.23843v1 Announce Type: new
Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compressio...
By Zizhong Wang, Jieying Wang, Zhao Zhang, Jiajia Li
arXiv:2607. 24331v1 Announce Type: new Abstract: As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually becomes a significant bottleneck as the context window continues to grow.
By Tan T. Nguyen, Quan V. Dang
arXiv:2602.08005v2 Announce Type: replace-cross
Abstract: Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computati...
By Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu
arXiv:2610.06927v1 Announce Type: cross
Abstract: The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free r...
By Sara Abdali, Jongwoo Ko, Pashmina Cameron
arXiv:2610.02953v1 Announce Type: new
Abstract: Long-context LLM serving is increasingly bottlenecked by KV-cache memory, especially in resource-constrained scenarios. Among existing KV-cache compres...
By Zihan Teng, Jiayu Zhao, Wentao Ren, Minhao Fan, Tianrui Ma, Song Chen, Weichen Liu
KV$^2$ is a query‑agnostic key‑value cache compression technique that selectively reconstructs only informative in‑context tokens using a lightweight proxy scorer before final eviction scoring. On benchmarks such as RULER, Needle‑in‑a‑Haystack, and LongBench, KV$^2$ outperforms baseline methods, especially under tight memory budgets, achieving higher scores with lower runtime and peak memory than full‑context reconstruction. The approach demonstrates that reusable KV‑cache compression can avoid reprocessing the entire prompt while maintaining quality.
By Johannes Wesch, Danni Liu, Jan Niehues