Hugging Face Blog

Faster Assisted Generation with Dynamic Speculation

Towards Data Science
Aug 24

Speculative Decoding on CPUs: Nearly 4x Faster Token Generation with DFlash

Speculative decoding leverages idle CPU resources to accelerate token generation without altering model outputs. In vLLM benchmarks, DFlash achieved a 3.92× increase in autoregressive throughput using Qwen3.5‑9B on an Intel Xeon 6 at a concurrency of 1. The article details the origins of this speedup, discusses acceptance metrics, and outlines factors that influence when speculation is beneficial.

By Ehssan Khan
Hugging Face Trending Papers
Jul 27

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.

arXiv Computation and Language
Sep 24

When Parallel Drafter Meets Parallel Speculative Decoding

The paper introduces DPara, a parallel speculative decoding framework that builds on DSpark-style parallel drafters. DPara eliminates the need for probabilistic guesses by precomputing draft representations for every acceptance boundary and using a lightweight autoregressive head to combine verification outcomes with these representations, enabling full parallelization of the backbone forward pass. Experiments on Qwen3-8B and Qwen3-14B across multiple benchmarks demonstrate average speedups of 3.21× and 3.52× over autoregressive decoding, outperforming existing serial and parallel speculative decoding methods.

By Fuliang Liu, Xue Li, Kun Qian, Zhibin Wang, Wanchun Dou, Wenyuan Yu, Chen Tian