arXiv Machine Learning

Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

arXiv:2608. 03893v1 Announce Type: new Abstract: Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch.

arXiv AI
Sep 2

CacheBridge: Efficient Cross-Model KV Cache Transfer

CacheBridge is a method for efficiently transferring key‑value (KV) caches between large language models (LLMs) in a multi‑model system. It replaces the full‑head mapping approach by matching each target KV head to a single source head, weighting reconstruction errors by causal attention sensitivity, and building weighted sufficient statistics with a fused GPU kernel. The technique achieves comparable or better accuracy to full‑head mapping while reducing mapper storage by up to eight‑fold, accelerating application by up to three‑times, and cutting construction time dramatically.

By Xingyu Qu, Siyuan Lu, Zhiyu Chen, Sheng Wang, Tao Lin
arXiv Computation and Language
Sep 10

Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

The paper addresses the challenge of long input contexts in Retrieval-Augmented Generation (RAG) systems, where concatenating many retrieved chunks increases prefill workload and time to first token (TTFT). It proposes a dual strategy: fine‑tuning the model to be aware of KV cache concatenation and selectively recomputing only part of the KV caches. Experiments on the RULER benchmark show that for a 124k‑token input, this combined method boosts the RULER score by 9.7 points over a baseline that recomputes caches only, while cutting TTFT by 80% compared with full attention.

By Fumihiko Tachibana, Daisuke Miyashita, Jun Deguchi