arXiv Machine Learning

When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding

arXiv:2606. 30265v1 Announce Type: new Abstract: Speculative decoding accelerates language model inference by using a fast drafter to propose candidate tokens that are then verified by a larger target model.

arXiv Computation and Language
Sep 7

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

The paper examines lossy verification techniques used in speculative decoding for large language models, showing that many methods can be grouped into truncation-based and collaborative verification categories. It analyzes how these approaches alter the decoding distribution, revealing that truncation-based methods can significantly degrade performance due to distributional distortion, while collaborative methods depend more on overshoot suppression and supervision quality than on simple interpolation between draft and target models. A diagnostic evaluation framework is introduced to assess these failure modes across curated benchmarks.

By Tianyu Wang, Yuxuan Zhou, Heng Li, Wenbin Wang, Zikai Xiao, Chunrui Zheng, Junyuan Shang
arXiv AI
Aug 26

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

ResiSpec is a framework that improves speculative decoding for large language models by reshaping the residual distribution during verification. It addresses the problem of residual drift, where rejected candidates cause the target distribution to diverge from the draft model’s predictions, rendering later candidates ineffective. By aligning the verification process with the draft model’s high‑confidence regions, ResiSpec prevents candidate obsolescence and achieves up to 1.92× speedup over existing multi‑candidate methods.

By Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye
arXiv Machine Learning
Jun 2

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

arXiv:2602. 23881v2 Announce Type: replace Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate tokens that are then verified in parallel by the target model.

By Alexander Samarin, Sergei Krutikov, Anton Shevtsov, Sergei Skvortsov, Filipp Fisin, Alexander Golubev
arXiv Machine Learning
5d ago

Mentored Decoding: Faster Inference meets Boosting

The paper introduces mentored decoding, a formal framework for lossy speculative decoding that can accelerate inference of autoregressive language models while potentially improving output quality. It connects this inference technique to boosting theory and extends it to all f‑divergences, revealing geometric insights for total variation, simple approximations tied to boosting compliance, and a divergence‑independent data structure enabling efficient optimal parameter queries and mentored distribution construction.

By Vivien Tran-Thien, Richard Nock