arXiv AI

HijackKV: New Threat in Position-Independent KV Cache Reuse

arXiv:2607. 19957v1 Announce Type: cross Abstract: Key-Value (KV) cache reduces inference latency in large language models (LLMs).

arXiv Machine Learning
Aug 28

Affix Cache for Diffusion Large Language Models

The paper introduces ACache, an affix-oriented cache reuse mechanism for Diffusion Large Language Models (DLLMs). ACache identifies a small set of critical affix tokens, called Anchor Tokens, and selectively recomputes their key-value states while reusing the rest of the affix cache. Experiments on Fast-dLLM and Nano-vLLM show that recomputing about 20% of affix tokens restores accuracy and can reduce recompute latency by up to 55.7% while improving throughput by up to 1.68×.

By Kaihua Liang, An Zhong, Xin Tan, Zafar Ayyub Qazi, Hong Xu, Jian Weng, Marco Canini
arXiv AI
Sep 10

Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces

The paper introduces an attack that reconstructs text generated by locally hosted large language models by monitoring CPU cache activity during detokenization. It uses Flush+Reload on shared tokenizer code to time decoding, then Prime+Probe to capture token‑dependent cache traces, followed by a clustering‑and‑language‑model pipeline to recover the output text. The method is evaluated across various datasets, hardware, inference frameworks, and model families, successfully retrieving semantically accurate outputs from real‑world local LLM deployments, including agentic systems.

By Roy Weiss, Benyamin Konstantinov, Eitam Sheetrit, Tomer Simon, Yisroel Mirsky
arXiv Computation and Language
Sep 10

Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

The paper addresses the challenge of long input contexts in Retrieval-Augmented Generation (RAG) systems, where concatenating many retrieved chunks increases prefill workload and time to first token (TTFT). It proposes a dual strategy: fine‑tuning the model to be aware of KV cache concatenation and selectively recomputing only part of the KV caches. Experiments on the RULER benchmark show that for a 124k‑token input, this combined method boosts the RULER score by 9.7 points over a baseline that recomputes caches only, while cutting TTFT by 80% compared with full attention.

By Fumihiko Tachibana, Daisuke Miyashita, Jun Deguchi
arXiv Machine Learning
5d ago

CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters

CacheReforge is a method for recovering stale key‑value (KV) caches in large language models when lightweight adapters evolve. It represents stale caches as layer‑wise mixed‑version objects and uses adapter anchors, sensitivity calibration, drift accumulation, and restart boundaries to decide between direct reuse, bounded recomputation, or full suffix recovery. Experiments on Qwen2.5 models with continual LoRA updates show a 92.4% reduction in mean KL divergence while only recomputing 5.44% of layers and cutting cache‑maintenance time by 93.2% compared to full prefill.

By Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen