arXiv Machine Learning

LumoTree: Path-Parallel Speculative Verification for Hybrid Language Models

LumoTree is a verifier designed for hybrid language models that ensures a single coherent continuation across recurrent, convolution, and attention states. It executes recurrent paths in parallel, reuses state tiles, and coordinates replay, history gathering, and cache remapping through a shared logical tree. Experiments demonstrate exact candidate‑selection parity and recurrent agreement within error bounds, and a Qwen3.8‑27B deployment achieved 25.63 pooled tokens/s on a single NVIDIA DGX Spark.

arXiv Machine Learning
Aug 4

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

arXiv:2608. 01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound.

By Li Wang, Yi Su, Xiabao Wu, Chiran You, Yongchao Liu, Zhan Qiu, Juelu Zhang, Jiajun Zheng, Fangxin Liu, Jie Zhang, Chen Tian, Chengying Huan
arXiv AI
Jul 7

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

arXiv:2607. 05147v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification.

By Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang
arXiv AI
2d ago

DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair

DRelay introduces a global draft context mechanism to improve prefix-aware parallel speculative decoding for large language models. By using a global reader to extract predictive information across the entire draft block and a causal selector to repair early token selection errors, DRelay extends the accepted prefix length and enhances decoding performance. Experiments on eight benchmarks show consistent gains over existing methods such as DFlash, Domino, and DSpark, with notable speedup improvements in SGLang serving.

By Zhuoyu Wang, Junnan Huang, Xinyu Chen
arXiv Machine Learning
Jul 17

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

arXiv:2607. 14952v1 Announce Type: new Abstract: A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment.

By Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, Irvine Lu, Nolan Ho, Lucian Li, Andrew Lei, Cleon Cheng, Steven Chiang, Yihang Zeng, Di Zhang, Rio Yang, Kaijie Chen, Andrew Chen, Pony Ma, Weizhong Zhang, Cheng Jin