arXiv AI

Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache

arXiv Machine Learning
5d ago

CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters

CacheReforge is a method for recovering stale key‑value (KV) caches in large language models when lightweight adapters evolve. It represents stale caches as layer‑wise mixed‑version objects and uses adapter anchors, sensitivity calibration, drift accumulation, and restart boundaries to decide between direct reuse, bounded recomputation, or full suffix recovery. Experiments on Qwen2.5 models with continual LoRA updates show a 92.4% reduction in mean KL divergence while only recomputing 5.44% of layers and cutting cache‑maintenance time by 93.2% compared to full prefill.

By Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen
arXiv AI
Jun 29

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

arXiv:2601. 16956v1 Announce Type: cross Abstract: The rapid growth of Large Transformer-based models, specifically Large Language Models (LLMs), now scaling to trillions of parameters, has necessitated training across thousands of GPUs using complex hybrid parallelism strategies (e.

By Avinash Maurya, M. Mustafa Rafique, Franck Cappello, Bogdan Nicolae
arXiv Computation and Language
Sep 23

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Flash-dLLM is a training‑free inference acceleration framework that improves the speed and memory efficiency of Diffusion Large Language Models (dLLMs). It tackles GPU memory I/O bottlenecks by introducing an I/O‑aware fused KV‑cache kernel and then employs a draft‑and‑verify decoding strategy that uses the dLLM itself as both drafter and verifier. Experiments on mathematical reasoning and code‑generation tasks show Flash‑dLLM outperforms existing acceleration methods, achieving up to 11.0× speedups over the Elastic‑Cache baseline.

By Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen