arXiv AI

ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover

arXiv:2608. 10545v1 Announce Type: cross Abstract: Edge LLMs must preserve inference continuity when a user hands over between edge nodes, requiring key-value (KV) cache transfer to the target node.

arXiv Computation and Language
Sep 7

Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference

The paper introduces a cache‑aware post‑training framework for Mixture‑of‑Experts (MoE) models that jointly adapts the MoE backbone and lightweight auxiliary cache routers while keeping the native Top‑K expert‑selection rule. Two modes are proposed: Temporal Router, which predicts same‑layer reuse and retains experts for future tokens, and Spatio‑Temporal Router, which adds a Spatio Router that refines the temporal cache using the causal predecessor’s hidden state. Experiments on Qwen3 and GPT‑OSS across GSM8K, MATH, and CommonsenseQA show that Temporal Router improves cache hit rates and reduces expert‑weight traffic, while Spatio‑Temporal Router achieves the best load‑adjusted efficiency, outperforming strong prefetching baselines.

By Zhenhe Wu, Yaping Jin, Qinghua Xing, Hang Zhou, Wei He, Xianjie Wu, Xianfu Cheng, Jian Yang, Hanting Chen
arXiv Machine Learning
Sep 1

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware

WiSP (Working‑Set Paging) is a routing‑aware expert pager that allows Mixture‑of‑Experts models to run on GPUs that cannot hold the entire expert pool by paging experts in and out of VRAM while preserving byte‑identical outputs. On a 24 GiB RTX 3090, WiSP doubles decode throughput compared to static offload when the model does not fit, and its companion policy MV‑WSA allocates VRAM between resident experts and KV cache based on marginal latency benefit, reducing end‑to‑end time by up to 1.19× without altering model outputs.

By Jiamu Zhang, Liang Wu, Mayank Darbari, Liangjie Hong
arXiv AI
Aug 20

Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study

The paper investigates whether training Mixture-of-Experts (MoE) routers can improve memory‑bandwidth locality on consumer GPUs. Using a new zero‑surgery telemetry tool, the authors measure that a large Qwen3‑235B model is bottlenecked by disk‑based expert access, and that an LRU cache can serve a majority of requests. They pre‑register experiments training 137 M‑parameter MoE models with locality‑aware losses, finding that while cache misses can drop up to 60 % (99 % static‑pin hit rate), every configuration fails to meet a strict 1 % perplexity threshold, indicating a tight coupling between cache efficiency and model quality.

By Shriniwas Ramesh Suram
arXiv Machine Learning
Aug 18

Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN

arXiv:2608. 16477v1 Announce Type: new Abstract: AI-RAN brings large language model (LLM) serving close to mobile users, but cellular handover can separate an active request from its inference state: the user attaches to a target base station (gNB) while the large and growing key-value (KV) cache remains at the source.

By Tianhang Ding, Jianchun Liu, Hongli Xu
Hugging Face Trending Papers
Jun 11

MiniPIC: Flexible Position-Independent Caching in <100LOC

Retrieval-augmented and agentic workloads repeatedly prefill recurring predictable structured inputs (which we call "spans") such as documents and code files. Yet, prefix caching in engines such as vLLM cannot reuse their KV entries unless they share identical prefixes with another request, while Position-Independent Caching (PIC) implementations within production-grade inference servers typically either require substantial server code changes or keep KV state outside the server, incurring host-to-device transfer overhead.

arXiv AI
Sep 2

CacheBridge: Efficient Cross-Model KV Cache Transfer

CacheBridge is a method for efficiently transferring key‑value (KV) caches between large language models (LLMs) in a multi‑model system. It replaces the full‑head mapping approach by matching each target KV head to a single source head, weighting reconstruction errors by causal attention sensitivity, and building weighted sufficient statistics with a fused GPU kernel. The technique achieves comparable or better accuracy to full‑head mapping while reducing mapper storage by up to eight‑fold, accelerating application by up to three‑times, and cutting construction time dramatically.

By Xingyu Qu, Siyuan Lu, Zhiyu Chen, Sheng Wang, Tao Lin