Faster Assisted Generation with Dynamic Speculation
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
Speculative decoding leverages idle CPU resources to accelerate token generation without altering model outputs. In vLLM benchmarks, DFlash achieved a 3.92× increase in autoregressive throughput using Qwen3.5‑9B on an Intel Xeon 6 at a concurrency of 1. The article details the origins of this speedup, discusses acceptance metrics, and outlines factors that influence when speculation is beneficial.
Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.
arXiv:2607. 24434v1 Announce Type: cross Abstract: Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory.
The paper introduces DPara, a parallel speculative decoding framework that builds on DSpark-style parallel drafters. DPara eliminates the need for probabilistic guesses by precomputing draft representations for every acceptance boundary and using a lightweight autoregressive head to combine verification outcomes with these representations, enabling full parallelization of the backbone forward pass. Experiments on Qwen3-8B and Qwen3-14B across multiple benchmarks demonstrate average speedups of 3.21× and 3.52× over autoregressive decoding, outperforming existing serial and parallel speculative decoding methods.