arXiv Computation and Language By Kaizhen Tan, Rong Gu, Mingyuan Li

CacheWeaver: Cache-Aware Evidence Ordering for Efficient Grounded RAG Inference

Read the original on arXiv Computation and Language →

CacheWeaver is a lightweight prompt‑layer technique that orders evidence for Retrieval‑Augmented Generation (RAG) to improve cache reuse in serving engines like vLLM. By maintaining a prefix tree of recently served evidence sequences and greedily placing the most reusable prefix first, it reduces median time‑to‑first‑token by 20‑33 % across three vLLM configurations without harming answer quality. The greedy policy achieves 97.5 % of the gain possible with oracle ordering, showing that most reusable prefix locality can be recovered with a simple scheduling layer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 6

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving

arXiv:2606. 05875v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves large language model (LLM) answer quality by grounding generation in external evidence, but processing retrieved contexts makes the prefill stage a dominant serving cost.

By Jianxin Yan, Wangze Ni, Zhenxin Li, Jiabao Jin, Zhitao Shen, Haoyang Li, Jia Zhu, Peng Cheng, Xuemin Lin, Lei Chen, Kui Ren
Hugging Face Trending Papers
Jun 4

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving

Retrieval-augmented generation (RAG) improves large language model (LLM) answer quality by grounding generation in external evidence, but processing retrieved contexts makes the prefill stage a dominant serving cost. RAG cache fusion reduces this cost by reusing precomputed key-value (KV) caches for retrieved chunks and selectively recomputing tokens under the current prompt.

arXiv Machine Learning
Sep 25

When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse

The paper investigates cache replacement strategies for large language model (LLM) prefix reuse, analyzing production traces from two companies and testing 14 eviction algorithms in both high-bandwidth memory (HBM) and large memory-pool environments. It finds that sophisticated policies designed for traditional caches offer little advantage over simple LRU, because prefix reuse is largely driven by the regular pacing of active sessions, making recency a strong predictor. The study also highlights new challenges such as heavy-tailed session footprints and variable miss costs, and proposes a compute-savings ratio along with two offline oracles to better quantify these effects, suggesting that effective prefix-cache management should combine recency with selective quick demotion, compute-aware partial eviction, and capacity-dependent granularity.

By Yiyu Liu, Minlan Yu, Juncheng Yang