arXiv AI By Ranran Haoran Zhang, Soumik Dey, Ashirbad Mishra, Hansi Wu, Binbin Li, Rui Zhang

Correctness Forensics for Batch Speculative Decoding: Diagnosing the Ragged Tensor Problem

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

Towards Data Science
Aug 24

Speculative Decoding on CPUs: Nearly 4x Faster Token Generation with DFlash

Speculative decoding leverages idle CPU resources to accelerate token generation without altering model outputs. In vLLM benchmarks, DFlash achieved a 3.92× increase in autoregressive throughput using Qwen3.5‑9B on an Intel Xeon 6 at a concurrency of 1. The article details the origins of this speedup, discusses acceptance metrics, and outlines factors that influence when speculation is beneficial.

By Ehssan Khan
arXiv AI
Jul 7

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

arXiv:2607. 05147v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification.

By Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang