arXiv Machine Learning By Qiao Hu, Yepeng Weng, Bo Zhang, Takehisa Yairi

RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding

Read the original on arXiv Machine Learning →

RheoSampling tackles the one‑hot dilemma in stochastic dynamic‑tree speculative decoding by decoupling token sampling from tree construction and verification. It injects a sampled token into deterministic top‑K slots, assigning distinct proxy probabilities for expansion and pruning while preserving the true sampling probability for verification, thereby achieving lossless, context‑aware top‑K construction with stochastic sampling. Experiments on various LLMs show higher acceptance rates and speedups compared to existing dynamic‑tree methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 28

TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

TreeGraft introduces a multi-drafter framework that combines drafters of varying costs to build a shared draft tree for tree-based speculative decoding. The stronger drafter rescues and rescoring candidates from the weaker drafter, while a lightweight scheduler decides when to invoke the stronger drafter to manage cost. Experiments on 10 model pairs and 6 benchmarks show TreeGraft improves over the best single-drafter strategy by an average of 15.1% and up to 26.6%.

By Jiaming Fan, Daming Cao, Canchen Huang, Jiale Fu, Jin Zhang, Junjie Gao, Kai Yang, Xiangzhong Luo, Xu Yang
Hugging Face Trending Papers
Aug 13

DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees

Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path.