arXiv Machine Learning By Yuchen Yang, Yifan Zhao, Anisha Dasgupta, Sasa Misailovic

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

Read the original on arXiv Machine Learning →

arXiv:2607. 16184v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) is a popular class of large language models (LLMs), offering high efficiency and accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
6d ago

DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory

DPS (Dual-Mode Precision LLM Serving) is a system that treats model‑weight memory as elastic by using a multi‑precision representation. Under normal load it serves the full‑accuracy model, but when KV‑cache pressure spikes it switches to a lower‑precision variant and reallocates unused weight memory for KV cache blocks. Built on Semi‑Unified Memory and implemented on top of vLLM, DPS boosts sustained throughput by 2.1–3.3× and effective pass@1 by up to +41 pp over static FP16 while maintaining FP16‑class accuracy.