arXiv AI By Daming Luo, Christy Liang, Junyu Xuan

How Query Visibility Changes KV-Cache Compression Rankings: A Matched-Budget Audit

Read the original on arXiv AI →

arXiv:2607. 11942v1 Announce Type: cross Abstract: KV-cache compression methods are predominantly evaluated with the query appended to the context before compression -- a query-aware protocol.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 27

Trust the Mass: Forced Weights in KV-Cache Eviction

The paper investigates KV‑cache eviction strategies for sparse‑attention models, showing that selecting the largest attention weights is nearly optimal—closing only a median 2–5 % of the gap to full attention. It further demonstrates that differences in performance between eviction methods largely stem from memory usage, with the new training‑free ContourKV allocator outperforming state‑of‑the‑art methods in most pairwise comparisons while matching their byte‑efficiency.

By Jack Shi, Jerry Gu
arXiv Machine Learning
Sep 4

VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch

VestigeKV is a new KV‑cache technique that uses a 64‑dimensional vestigial branch—originally a RoPE component repurposed during NoPE training—as a query‑independent eviction signal. By reading only 11 % of each cache row, the method partitions the cache into an attended tier (top‑m rows) and an archive tier (all other rows), which is GPU‑resident and never deleted. The approach achieves near‑perfect retrieval (1.00 at 8×, 0.92 at 32×) without any training, quantization, or changes to weights or kernels.

By WenJie Fan