arXiv Machine Learning By Yiyu Liu, Minlan Yu, Juncheng Yang

When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse

Read the original on arXiv Machine Learning →

The paper investigates cache replacement strategies for large language model (LLM) prefix reuse, analyzing production traces from two companies and testing 14 eviction algorithms in both high-bandwidth memory (HBM) and large memory-pool environments. It finds that sophisticated policies designed for traditional caches offer little advantage over simple LRU, because prefix reuse is largely driven by the regular pacing of active sessions, making recency a strong predictor. The study also highlights new challenges such as heavy-tailed session footprints and variable miss costs, and proposes a compute-savings ratio along with two offline oracles to better quantify these effects, suggesting that effective prefix-cache management should combine recency with selective quick demotion, compute-aware partial eviction, and capacity-dependent granularity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 18

PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving

PrefixBench-H100 is a reproducible benchmark that evaluates how reusing prompt prefixes affects LLM serving performance on NVIDIA H100 GPUs. It tests two popular runtimes (vLLM and TensorRT-LLM) across varied workloads, measuring metrics such as time-to-first-token, latency, throughput, cache hits, and GPU memory usage. The study identifies when prefix reuse significantly reduces first‑token latency and when cache pressure diminishes those gains, noting that cache effectiveness is largely unaffected by concurrency or output length, while differences arise mainly in scheduling.

By Omkar Shewale, Deepak Kumar, Divakar Kumar Yadav
arXiv Machine Learning
Sep 7

Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving

The paper investigates how prefix caching, a default optimization in open‑source LLM serving stacks, affects reproducibility when combined with weight quantization. Experiments on an eighty‑episode multi‑turn agentic tool‑use workload show that enabling the cache causes the agent’s trajectory to change in 36.2 % of episodes at 16‑bit precision and 75.0 % at 4‑bit precision, while disabling the cache yields perfectly reproducible runs. The study identifies specific cache‑related settings that drive run‑to‑run divergence and demonstrates that cached serving is deterministic only when the cache state is preserved, which is not the case in typical deployments.

By Aditi Patodiya
arXiv Computation and Language
Aug 27

TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving

TOPAS is a Task‑Oriented Prefix‑Aware Scheduler designed for multi‑agent large language model serving. It jointly decides which agent prefixes to retain in a shared key‑value cache and which requests to schedule, balancing the reduction of each task’s longest remaining service path against the benefit of downstream prefix reuse while accounting for movement and preemption costs. Experiments on synthetic DAGs and MetaGPT software‑development workflows show that TOPAS can reduce mean and p99 job completion times by up to 39.8%/49.4% and 22.0%/26.6% respectively compared to the best baselines.

By Hongqiu Ni, Han Tian, Chi Zhang, Guopeng Li, Haisheng Tan