arXiv AI By Johannes Wesch, Danni Liu, Jan Niehues

KV$^2$: A Self-Refining KV Cache

Read the original on arXiv AI →

KV$^2$ is a query‑agnostic key‑value cache compression technique that selectively reconstructs only informative in‑context tokens using a lightweight proxy scorer before final eviction scoring. On benchmarks such as RULER, Needle‑in‑a‑Haystack, and LongBench, KV$^2$ outperforms baseline methods, especially under tight memory budgets, achieving higher scores with lower runtime and peak memory than full‑context reconstruction. The approach demonstrates that reusable KV‑cache compression can avoid reprocessing the entire prompt while maintaining quality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
5d ago

EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction

EchoPress is a training‑free method for pruning key‑value caches in large language models. It approximates the reconstruction attention used by KVzip by leveraging queries and keys from standard prefill, reconstructing only the first chunk to calibrate importance scores for the rest of the context. Experiments on LongBench and RULER with Qwen3‑8B and Llama‑3.1‑8B‑Instruct show that EchoPress matches KVzip’s task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by 1.7–19.6× and total prefill time by up to 2.9×.

By Jiawei Lin, Saibo Geng, Thomas Bourgeat
arXiv AI
Jun 24

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

arXiv:2606. 24467v1 Announce Type: new Abstract: Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware.

By Xiaolin Lin, Jingcun Wang, Olga Kondrateva, Yiyu Shi, Bing Li, Grace Li Zhang