arXiv Machine Learning By Yanxiao Liu, Sicheng Wan, Zhan Gao, Deniz G\"und\"uz

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

Read the original on arXiv Machine Learning →

The paper introduces a new multi-draft speculative sampling algorithm that uses Poisson processes to improve inference efficiency and output provenance for large language models. It achieves strong sampling efficiency while allowing an unbiased watermark to be embedded without reducing speculative acceptance. The method relies on an exact list‑coupling‑without‑communication scheme, giving a drafter‑invariant property that benefits both sampling and watermarking, and the authors experimentally confirm its effectiveness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 26

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

ResiSpec is a framework that improves speculative decoding for large language models by reshaping the residual distribution during verification. It addresses the problem of residual drift, where rejected candidates cause the target distribution to diverge from the draft model’s predictions, rendering later candidates ineffective. By aligning the verification process with the draft model’s high‑confidence regions, ResiSpec prevents candidate obsolescence and achieves up to 1.92× speedup over existing multi‑candidate methods.

By Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye