DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 05147v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification.
The paper introduces DPara, a parallel speculative decoding framework that builds on DSpark-style parallel drafters. DPara eliminates the need for probabilistic guesses by precomputing draft representations for every acceptance boundary and using a lightweight autoregressive head to combine verification outcomes with these representations, enabling full parallelization of the backbone forward pass. Experiments on Qwen3-8B and Qwen3-14B across multiple benchmarks demonstrate average speedups of 3.21× and 3.52× over autoregressive decoding, outperforming existing serial and parallel speculative decoding methods.
Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.
arXiv:2609.24698v1 Announce Type: new Abstract: Repeated execution of the target model during autoregressive decoding is a major source of LLM inference latency. Unlike linear speculation, which foll...
arXiv:2609.37029v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context...
arXiv:2512. 22420v5 Announce Type: replace-cross Abstract: Speculative decoding (SD) accelerates LLM inference by verifying draft tokens in parallel.