The paper introduces AdaptiveSpec, a training‑free speculative decoding method that simultaneously adapts the per‑step verification rule and the draft‑tree shape using signals generated during decoding. It replaces the fixed token‑match rule with a margin‑based threshold and adjusts tree depth, width, and node count based on draft confidence and recent acceptance history, allowing the total draft count to vary. Experiments on SGLang show up to 56% throughput gains over EAGLE‑3 while maintaining 93% of lossless task accuracy on GSM8K, MATH‑500, and HumanEval across three models.
By Oszk\'ar Urb\'an, Young D. Kwon, Stylianos I. Venieris, Cecilia Mascolo
arXiv:2609.24698v1 Announce Type: new
Abstract: Repeated execution of the target model during autoregressive decoding is a major source of LLM inference latency. Unlike linear speculation, which foll...
By Changxu Liu, Zhaogeng Li
The paper introduces DPara, a parallel speculative decoding framework that builds on DSpark-style parallel drafters. DPara eliminates the need for probabilistic guesses by precomputing draft representations for every acceptance boundary and using a lightweight autoregressive head to combine verification outcomes with these representations, enabling full parallelization of the backbone forward pass. Experiments on Qwen3-8B and Qwen3-14B across multiple benchmarks demonstrate average speedups of 3.21× and 3.52× over autoregressive decoding, outperforming existing serial and parallel speculative decoding methods.
By Fuliang Liu, Xue Li, Kun Qian, Zhibin Wang, Wanchun Dou, Wenyuan Yu, Chen Tian
arXiv:2607. 12422v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive language model inference by using a cheap drafter to propose multiple future tokens and a target model to verify them.
By Abdurrahman Javat, Allan Kazakov
arXiv:2608.30427v1 Announce Type: cross
Abstract: Speculative decoding speeds up generation with an efficient draft model (drafter) that proposes tokens for a target model to verify in one pass, pres...
By Ephrem Wu
arXiv:2609.14717v1 Announce Type: cross
Abstract: Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to accept seve...
By Jahyun Koo, Sunghyeon Woo, Jaeeun Kil, Jeongtae Lee, Sungjae Lee, Kyomin Jung, Minsub Kim