TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding
arXiv:2606. 00487v1 Announce Type: new Abstract: Using a diffusion model for parallel drafting is a promising approach for speculative decoding.
The paper introduces DPara, a parallel speculative decoding framework that builds on DSpark-style parallel drafters. DPara eliminates the need for probabilistic guesses by precomputing draft representations for every acceptance boundary and using a lightweight autoregressive head to combine verification outcomes with these representations, enabling full parallelization of the backbone forward pass. Experiments on Qwen3-8B and Qwen3-14B across multiple benchmarks demonstrate average speedups of 3.21× and 3.52× over autoregressive decoding, outperforming existing serial and parallel speculative decoding methods.
arXiv:2606. 00487v1 Announce Type: new Abstract: Using a diffusion model for parallel drafting is a promising approach for speculative decoding.
arXiv:2607. 05147v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification.
Speculative decoding accelerates autoregressive language model inference by using a cheap drafter to propose multiple future tokens and a target model to verify them. A common design goal is therefore to improve draft quality while reducing auxiliary parameters and systems overhead.
arXiv:2607. 12422v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive language model inference by using a cheap drafter to propose multiple future tokens and a target model to verify them.
arXiv:2606. 03819v1 Announce Type: new Abstract: One-shot block drafters for speculative decoding generate the full draft in a single forward pass, achieving strong throughput by eliminating sequential token generation.
arXiv:2606. 04446v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive large language model inference by drafting multiple tokens and verifying them in a single target-model forward pass.
arXiv:2603. 18016v2 Announce Type: replace-cross Abstract: Speculative decoding (SD) accelerates large language model inference by using a smaller draft model to propose draft tokens that are subsequently verified by a larger target model.
arXiv:2608. 08721v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency critically determined by the speculative length selected at each decoding round.
arXiv:2609.37029v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context...
arXiv:2609.38510v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusi...
The paper introduces TSS, a target-side sparsification framework that selectively skips layers in a target verifier during speculative decoding for domain-specific large language models. By exploring multi-layer skip configurations with an acceptance- and metric-aware breadth search, TSS reduces verification cost, increases draft acceptance, and can even improve downstream task performance without retraining. Experiments on Spec-Bench demonstrate consistent gains across domains and model scales, notably boosting translation throughput by 1.68× and improving BLEU scores significantly.
ResiSpec is a framework that improves speculative decoding for large language models by reshaping the residual distribution during verification. It addresses the problem of residual drift, where rejected candidates cause the target distribution to diverge from the draft model’s predictions, rendering later candidates ineffective. By aligning the verification process with the draft model’s high‑confidence regions, ResiSpec prevents candidate obsolescence and achieves up to 1.92× speedup over existing multi‑candidate methods.