arXiv AI By Chao Fei, Panos Kalnis

PolyKV: Heterogeneous Retention and Allocation for KV Cache Compression

Read the original on arXiv AI →

arXiv:2606. 15157v1 Announce Type: cross Abstract: KV cache compression is essential for reducing the memory cost of long-context large language model inference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

Exploring a Layer-Wise Design Space for KV Cache Eviction

The paper investigates whether key‑value (KV) cache eviction strategies should vary across Transformer layers. By combining existing eviction methods in different layer configurations and profiling their performance, the authors find that heterogeneous, layer‑wise routing consistently outperforms homogeneous policies on LongBench tasks. Even with a fixed set of methods, the placement of each method strongly influences overall quality, and a single well‑chosen route surpasses all nine standalone baselines across multiple cache budgets.

By Chao Fei, Kaihua Liang, Hanzhi Hu, Hongcheng Guo, Jian Weng, Marco Canini, Panos Kalnis
arXiv Machine Learning
Jul 3

The risk of KV cache compression

arXiv:2607. 01520v1 Announce Type: new Abstract: Transformer inference on long sequences is expensive because softmax attention repeatedly reads from a large KV cache.

By Lukas Haverbeck, Carmen Amo Alonso, Andres Felipe Posada-Moreno, Sebastian Trimpe, Marco Pavone
arXiv Machine Learning
Sep 17

GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference

GroupKV is a lightweight hierarchical KV cache management system designed for long‑context diffusion large language model (dLLM) inference. It partitions the context into contiguous groups and uses coarse‑to‑fine sparse selection, cross‑layer consistency for predictive prefetching, and a staleness correction mechanism to keep the cache coherent amid dynamic KV updates. The approach also incorporates streaming prefill to lower peak memory usage, achieving up to 48× longer serviceable context, 3.73× faster inference in offload‑based settings, and competitive task accuracy.

By Jinhao Wang, Zhexin Hu, Kangjie Zhou, Xin Zhou, Fangfang Liu