Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 20537v1 Announce Type: new Abstract: Mainstream LLM serving systems reuse prefix work mainly through paged or radix key-value (KV) caches.
CacheReforge is a method for recovering stale key‑value (KV) caches in large language models when lightweight adapters evolve. It represents stale caches as layer‑wise mixed‑version objects and uses adapter anchors, sensitivity calibration, drift accumulation, and restart boundaries to decide between direct reuse, bounded recomputation, or full suffix recovery. Experiments on Qwen2.5 models with continual LoRA updates show a 92.4% reduction in mean KL divergence while only recomputing 5.44% of layers and cutting cache‑maintenance time by 93.2% compared to full prefill.
arXiv:2607. 28495v1 Announce Type: new Abstract: Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that originally reached the prefix.
arXiv:2601. 16956v1 Announce Type: cross Abstract: The rapid growth of Large Transformer-based models, specifically Large Language Models (LLMs), now scaling to trillions of parameters, has necessitated training across thousands of GPUs using complex hybrid parallelism strategies (e.
arXiv:2607. 27539v2 Announce Type: replace Abstract: Exact deletion from persistent language-model memory depends on whether a record's effect remains addressable after later computation.
arXiv:2604. 26968v2 Announce Type: replace-cross Abstract: Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference serving.