arXiv Machine Learning By Jungseob Lee, Sugyeong Eo

CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters

Read the original on arXiv Machine Learning →

CAST (Cost‑Aware Speculative Trees) is a method that improves speculative decoding for large language models by packing multiple drafted token candidates into a tree and verifying the entire tree in a single target‑model pass, rather than only the top‑scoring chain. The tree width is adaptively chosen based on a latency measurement, ensuring that each added candidate’s expected gain outweighs its verification cost. Experiments across five domains, three GPU generations, and two model families show that CAST can be up to 43 % faster than the standard chain, while preserving the target model’s output distribution under both greedy and sampled decoding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 28

TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

TreeGraft introduces a multi-drafter framework that combines drafters of varying costs to build a shared draft tree for tree-based speculative decoding. The stronger drafter rescues and rescoring candidates from the weaker drafter, while a lightweight scheduler decides when to invoke the stronger drafter to manage cost. Experiments on 10 model pairs and 6 benchmarks show TreeGraft improves over the best single-drafter strategy by an average of 15.1% and up to 26.6%.

By Jiaming Fan, Daming Cao, Canchen Huang, Jiale Fu, Jin Zhang, Junjie Gao, Kai Yang, Xiangzhong Luo, Xu Yang