Hugging Face Trending Papers

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse under tight budgets. Existing methods primarily improve which original KV pairs are retained.

arXiv Computation and Language
Sep 7

Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation

The paper investigates how to recover language model quality lost when using low‑bit key–value caches for autoregressive decoding. By keeping the quantizer fixed and distilling the full‑precision cache behavior into low‑rank Q/K/V projection updates, the authors demonstrate that 4‑bit affine‑cache adapters recover roughly 54 % of the perplexity gap on TinyLlama‑1.1B and 76 % on Gemma‑4‑12B, while preserving most long‑context retrieval. Even a 2‑bit rank–token sweep can dramatically reduce TinyLlama’s perplexity, though it only partially restores retrieval performance.

By Seifeldin Abdellatif
arXiv Computation and Language
Sep 10

Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

The paper addresses the challenge of long input contexts in Retrieval-Augmented Generation (RAG) systems, where concatenating many retrieved chunks increases prefill workload and time to first token (TTFT). It proposes a dual strategy: fine‑tuning the model to be aware of KV cache concatenation and selectively recomputing only part of the KV caches. Experiments on the RULER benchmark show that for a 124k‑token input, this combined method boosts the RULER score by 9.7 points over a baseline that recomputes caches only, while cutting TTFT by 80% compared with full attention.

By Fumihiko Tachibana, Daisuke Miyashita, Jun Deguchi
arXiv AI
Jun 24

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

arXiv:2606. 24467v1 Announce Type: new Abstract: Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware.

By Xiaolin Lin, Jingcun Wang, Olga Kondrateva, Yiyu Shi, Bing Li, Grace Li Zhang
arXiv Machine Learning
1d ago

EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction

EchoPress is a training‑free method for pruning key‑value caches in large language models. It approximates the reconstruction attention used by KVzip by leveraging queries and keys from standard prefill, reconstructing only the first chunk to calibrate importance scores for the rest of the context. Experiments on LongBench and RULER with Qwen3‑8B and Llama‑3.1‑8B‑Instruct show that EchoPress matches KVzip’s task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by 1.7–19.6× and total prefill time by up to 2.9×.

By Jiawei Lin, Saibo Geng, Thomas Bourgeat