arXiv AI

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving

arXiv:2607. 02574v1 Announce Type: cross Abstract: The key-value (KV) cache has become a first-order memory object in LLM serving rather than a temporary per-request tensor.

arXiv AI
Jul 28

PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving

arXiv:2607. 22648v1 Announce Type: new Abstract: Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form reusable KV cache prefixes, which are the main technique used by state of art approaches to accelerate inferences.

By Meghana Maghyastha, Robert Underwood, Randal Burns, Bogdan Nicolae
Hugging Face Trending Papers
6d ago

DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory

DPS (Dual-Mode Precision LLM Serving) is a system that treats model‑weight memory as elastic by using a multi‑precision representation. Under normal load it serves the full‑accuracy model, but when KV‑cache pressure spikes it switches to a lower‑precision variant and reallocates unused weight memory for KV cache blocks. Built on Semi‑Unified Memory and implemented on top of vLLM, DPS boosts sustained throughput by 2.1–3.3× and effective pass@1 by up to +41 pp over static FP16 while maintaining FP16‑class accuracy.

arXiv AI
Jul 1

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving

arXiv:2605. 09735v2 Announce Type: replace-cross Abstract: Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular KV-cache behavior: request lengths differ, EOS events arrive asynchronously, and logical histories fragment over time.

By Zhiqing Zhong, Zhijing Ye, Jian Zhang, Weijian Zheng, Bolun Sun, Xiaodong Yu
arXiv AI
Sep 24

Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management

The paper introduces LM‑CXD, a CXL‑SSD design tailored for large language model (LLM) prefix caching. By aligning KV chunk management between the serving engine and the storage device, exposing NAND-to‑DRAM progress, and using device DRAM as a GPU‑accessible buffer, LM‑CXD reduces time‑to‑first‑token (TTFT) by up to 4× compared to a stock CXL‑SSD and brings performance within 1.5× of local DRAM across five LLM models. The approach also incorporates windowed prefetching and layer‑wise KV movement to hide NAND latency under limited device DRAM.

By Hyunsun Chung, Taewan Noh, Minji Kim, Joo-Young Hwang, Hong-Yeon Kim, Youngjae Kim