Hugging Face Trending Papers

DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory

Read the original on Hugging Face Trending Papers →

DPS (Dual-Mode Precision LLM Serving) is a system that treats model‑weight memory as elastic by using a multi‑precision representation. Under normal load it serves the full‑accuracy model, but when KV‑cache pressure spikes it switches to a lower‑precision variant and reallocates unused weight memory for KV cache blocks. Built on Semi‑Unified Memory and implemented on top of vLLM, DPS boosts sustained throughput by 2.1–3.3× and effective pass@1 by up to +41 pp over static FP16 while maintaining FP16‑class accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.