Sebastian Raschka

Understanding and Coding the KV Cache in LLMs from Scratch

KV caches are one of the most critical techniques for efficient inference in LLMs in production.

arXiv Machine Learning
Sep 29

Evaluating the accuracy of KV cache reuse techniques

The paper discusses position‑independent KV cache reuse, a technique designed to cut latency in retrieval‑augmented generation by reusing chunk‑level KV caches across prompts. It argues that current evaluation methods overstate the accuracy of such reuse because they do not accurately capture the loss of accuracy, and that existing datasets lack the necessary reuse dynamics for thorough testing. To remedy this, the authors propose a new evaluation methodology that unambiguously measures accuracy loss and introduce Boxoffice, a tool that programmatically creates datasets with challenging KV cache reuse patterns.

By Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona
arXiv AI
Jul 28

PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving

arXiv:2607. 22648v1 Announce Type: new Abstract: Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form reusable KV cache prefixes, which are the main technique used by state of art approaches to accelerate inferences.

By Meghana Maghyastha, Robert Underwood, Randal Burns, Bogdan Nicolae
arXiv Machine Learning
Aug 5

Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

arXiv:2608. 03893v1 Announce Type: new Abstract: Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch.

By Taekyung Heo, Rasoul Shafipour, Ritchie Zhao, Maximilian Golub, Mohammad Mahdi Kamani, Ritika Borkar, Makesh Tarun Chandran, Pantea Zardoshti, Bita Darvish Rouhani
arXiv AI
Jul 8

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

arXiv:2607. 06519v1 Announce Type: new Abstract: Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the layer-specific evidence needed for retrieval and multi-step reasoning.

By Anna C\'ordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jes\'us Olivera