arXiv AI By Chao Fei, Kaihua Liang, Hanzhi Hu, Hongcheng Guo, Jian Weng, Marco Canini, Panos Kalnis

Exploring a Layer-Wise Design Space for KV Cache Eviction

Read the original on arXiv AI →

The paper investigates whether key‑value (KV) cache eviction strategies should vary across Transformer layers. By combining existing eviction methods in different layer configurations and profiling their performance, the authors find that heterogeneous, layer‑wise routing consistently outperforms homogeneous policies on LongBench tasks. Even with a fixed set of methods, the placement of each method strongly influences overall quality, and a single well‑chosen route surpasses all nine standalone baselines across multiple cache budgets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 3

The risk of KV cache compression

arXiv:2607. 01520v1 Announce Type: new Abstract: Transformer inference on long sequences is expensive because softmax attention repeatedly reads from a large KV cache.

By Lukas Haverbeck, Carmen Amo Alonso, Andres Felipe Posada-Moreno, Sebastian Trimpe, Marco Pavone