Hugging Face Trending Papers

Accelerating Speculative Diffusions via Block Verification

Speculative decoding speeds up LLM inference by using a draft model to generate tokens, with an acceptance-rejection scheme that ensures that the output matches the target distribution. Adapting this to continuous diffusions is difficult because speculative sampling requires drawing from a residual distribution.

arXiv AI
Aug 26

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

ResiSpec is a framework that improves speculative decoding for large language models by reshaping the residual distribution during verification. It addresses the problem of residual drift, where rejected candidates cause the target distribution to diverge from the draft model’s predictions, rendering later candidates ineffective. By aligning the verification process with the draft model’s high‑confidence regions, ResiSpec prevents candidate obsolescence and achieves up to 1.92× speedup over existing multi‑candidate methods.

By Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye
arXiv AI
Sep 17

Accelerating Diffusion Sampling via Speculative Draft Trees

The paper introduces "draft trees" to accelerate diffusion model sampling by allowing non‑linear lookahead drafts, thereby increasing acceptance rates per expensive target evaluation. It connects speculative sampling to relative entropy coding, adopts greedy rejection sampling as the draft‑target coupling, and demonstrates up to 8.3% speed‑up over reflection coupling baselines in experiments.

By Marcello Bullo, Yanxiao Liu, \"Oyk\"u S{\i}la G\"uner, Arpan Mukherjee, Deniz G\"und\"uz
arXiv Machine Learning
Sep 21

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

The paper introduces a new multi-draft speculative sampling algorithm that uses Poisson processes to improve inference efficiency and output provenance for large language models. It achieves strong sampling efficiency while allowing an unbiased watermark to be embedded without reducing speculative acceptance. The method relies on an exact list‑coupling‑without‑communication scheme, giving a drafter‑invariant property that benefits both sampling and watermarking, and the authors experimentally confirm its effectiveness.

By Yanxiao Liu, Sicheng Wan, Zhan Gao, Deniz G\"und\"uz
arXiv Machine Learning
Jul 22

AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

arXiv:2607. 19223v1 Announce Type: new Abstract: Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified in parallel by the target model, has become a prevalent paradigm for accelerating large language model inference.

By Yu-Yang Qian, Hao-Cong Wu, Chen Chen, Jiacheng Sun, Zhenhua Dong, Peng Zhao, Zhi-Hua Zhou