arXiv AI By Zhuoyu Wang, Junnan Huang, Xinyu Chen

DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair

Read the original on arXiv AI →

DRelay introduces a global draft context mechanism to improve prefix-aware parallel speculative decoding for large language models. By using a global reader to extract predictive information across the entire draft block and a causal selector to repair early token selection errors, DRelay extends the accepted prefix length and enhances decoding performance. Experiments on eight benchmarks show consistent gains over existing methods such as DFlash, Domino, and DSpark, with notable speedup improvements in SGLang serving.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
2d ago

DEdit: Iterative Draft Editing for Speculative Decoding

arXiv:2609.38510v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusi...

By Longxuan Yu, Bingsen Chen, Peng Shi, Dongkyu Lee, Yi Xiang, Hideo Kobayashi, Sheng Zhang, Shuaichen Chang, Xing Niu, Zhuoyan Xu, Greg Ver Steeg, Jiarong Jiang
arXiv Computation and Language
Sep 4

Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding

The paper introduces AdaptiveSpec, a training‑free speculative decoding method that simultaneously adapts the per‑step verification rule and the draft‑tree shape using signals generated during decoding. It replaces the fixed token‑match rule with a margin‑based threshold and adjusts tree depth, width, and node count based on draft confidence and recent acceptance history, allowing the total draft count to vary. Experiments on SGLang show up to 56% throughput gains over EAGLE‑3 while maintaining 93% of lossless task accuracy on GSM8K, MATH‑500, and HumanEval across three models.

By Oszk\'ar Urb\'an, Young D. Kwon, Stylianos I. Venieris, Cecilia Mascolo
Hugging Face Trending Papers
Aug 13

DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees

Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path.