arXiv AI By Lifu Wang, Pan Zhou

SMART: When is it Actually Worth Expanding a Speculative Tree?

Read the original on arXiv AI →

arXiv:2604. 09731v2 Announce Type: replace-cross Abstract: Tree-based speculative decoding accelerates autoregressive generation by verifying a branching tree of draft tokens in a single target-model forward pass.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 4

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

arXiv:2608. 01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound.

By Li Wang, Yi Su, Xiabao Wu, Chiran You, Yongchao Liu, Zhan Qiu, Juelu Zhang, Jiajun Zheng, Fangxin Liu, Jie Zhang, Chen Tian, Chengying Huan
arXiv Machine Learning
1d ago

CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters

CAST (Cost‑Aware Speculative Trees) is a method that improves speculative decoding for large language models by packing multiple drafted token candidates into a tree and verifying the entire tree in a single target‑model pass, rather than only the top‑scoring chain. The tree width is adaptively chosen based on a latency measurement, ensuring that each added candidate’s expected gain outweighs its verification cost. Experiments across five domains, three GPU generations, and two model families show that CAST can be up to 43 % faster than the standard chain, while preserving the target model’s output distribution under both greedy and sampled decoding.

By Jungseob Lee, Sugyeong Eo
arXiv Computation and Language
Aug 28

TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

TreeGraft introduces a multi-drafter framework that combines drafters of varying costs to build a shared draft tree for tree-based speculative decoding. The stronger drafter rescues and rescoring candidates from the weaker drafter, while a lightweight scheduler decides when to invoke the stronger drafter to manage cost. Experiments on 10 model pairs and 6 benchmarks show TreeGraft improves over the best single-drafter strategy by an average of 15.1% and up to 26.6%.

By Jiaming Fan, Daming Cao, Canchen Huang, Jiale Fu, Jin Zhang, Junjie Gao, Kai Yang, Xiangzhong Luo, Xu Yang