arXiv AI By Zhuoren Ye, Tianyu Wo, Dinghao Xue, Mingming Zhang, Yuchen Teng, Chunming Hu, Renyu Yang

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

Read the original on arXiv AI →

arXiv:2606. 24506v1 Announce Type: cross Abstract: Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
6d ago

DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory

DPS (Dual-Mode Precision LLM Serving) is a system that treats model‑weight memory as elastic by using a multi‑precision representation. Under normal load it serves the full‑accuracy model, but when KV‑cache pressure spikes it switches to a lower‑precision variant and reallocates unused weight memory for KV cache blocks. Built on Semi‑Unified Memory and implemented on top of vLLM, DPS boosts sustained throughput by 2.1–3.3× and effective pass@1 by up to +41 pp over static FP16 while maintaining FP16‑class accuracy.