arXiv AI By Zipeng Gao, Zhi Zheng, Qingrong Xia, Junda Lin, Ziwei Zhao, Tong Xu, Zhefeng Wang, Enhong Chen

Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting

Read the original on arXiv AI →

arXiv:2607. 10661v1 Announce Type: cross Abstract: Speculative decoding has significantly accelerated Large Language Model (LLM) inference by alleviating memory-bound bottlenecks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 21

AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified in parallel by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work such as DFlash further boosts drafting efficiency by leveraging diffusion drafters, whose parallel denoising mechanism enables draft generation in a single forward pass.

arXiv Computation and Language
Sep 18

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

The paper introduces SwitchSD, an adaptive framework that controls speculative decoding by distinguishing genuine copy intent from accidental repetitions using lightweight probes on a model’s internal representations. SwitchSD dynamically switches between neural drafting and context-based copying, achieving up to 15% throughput gains over existing baselines such as EAGLE3. The approach turns copying from a noisy heuristic into a principled, model-aware decoding regime.

By Roy Eisenstadt, Ido Cohen, Edo Cohen-Karlik, Lior Wolf, Itamar Zimerman