arXiv AI

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation

arXiv:2602. 08686v3 Announce Type: replace-cross Abstract: Prefill-only KV compression freezes a token subset at the end of prefill and decodes from it without further eviction.

arXiv Machine Learning
Aug 27

Trust the Mass: Forced Weights in KV-Cache Eviction

The paper investigates KV‑cache eviction strategies for sparse‑attention models, showing that selecting the largest attention weights is nearly optimal—closing only a median 2–5 % of the gap to full attention. It further demonstrates that differences in performance between eviction methods largely stem from memory usage, with the new training‑free ContourKV allocator outperforming state‑of‑the‑art methods in most pairwise comparisons while matching their byte‑efficiency.

By Jack Shi, Jerry Gu
arXiv Machine Learning
Sep 22

PAGE: Partition-Aware Gated KV-Cache Eviction

PAGE is a partition‑aware gated KV‑cache eviction method that reframes eviction as a per‑input admission decision. It uses a single label‑free scalar— the early‑to‑late drop in pairwise top‑k head agreement—to classify inputs into a capacity‑bound class (where eviction is catastrophic) and a dilution‑prone class (where eviction is safe or beneficial). By thresholding this drop, PAGE applies a base evictor only when necessary, reducing the harm rate in the capacity‑bound regime from 0.75 to 0.026 and achieving a 29× improvement across four models and benchmarks without retraining the evictor.

By Pankaj Kumar, Subhankar Mishra
arXiv Machine Learning
Aug 31

Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation

The paper introduces PASK (Parser-Aware Structural KV Persistence), a method that leverages parser transitions to inform key‑value (KV) persistence decisions in structured generation tasks. By aligning KV compression with task‑level structured risk, PASK sets protection floors based on error sensitivity and allocates remaining KV capacity using attention‑output distortion, producing a lightweight, structure‑conditioned lookup policy. In experiments on Qwen3‑4B, PASK achieves a 17.39‑point accuracy gain over the best compressed baseline, while delivering up to 2.2× higher throughput, 3.3× lower TPOT, and 0.53× the peak GPU memory of full KV.

By Linze Wu, Xinrui Chen