arXiv AI

VIA-SD: Verification via Intra-Model Routing for Speculative Decoding

arXiv:2606. 12243v1 Announce Type: cross Abstract: Speculative decoding (SD) addresses the high inference costs of LLMs by having lightweight drafters generate candidates for large verifiers to validate in parallel.

arXiv Computation and Language
Sep 23

TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models

The paper introduces TSS, a target-side sparsification framework that selectively skips layers in a target verifier during speculative decoding for domain-specific large language models. By exploring multi-layer skip configurations with an acceptance- and metric-aware breadth search, TSS reduces verification cost, increases draft acceptance, and can even improve downstream task performance without retraining. Experiments on Spec-Bench demonstrate consistent gains across domains and model scales, notably boosting translation throughput by 1.68× and improving BLEU scores significantly.

By Haibo Hu, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue
arXiv AI
Jul 7

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

arXiv:2607. 05147v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification.

By Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang
arXiv AI
Aug 26

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

ResiSpec is a framework that improves speculative decoding for large language models by reshaping the residual distribution during verification. It addresses the problem of residual drift, where rejected candidates cause the target distribution to diverge from the draft model’s predictions, rendering later candidates ineffective. By aligning the verification process with the draft model’s high‑confidence regions, ResiSpec prevents candidate obsolescence and achieves up to 1.92× speedup over existing multi‑candidate methods.

By Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye
arXiv AI
Sep 16

FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

FlexEE is an early‑exiting framework designed for large language model inference that is constrained by computation and memory, particularly in offloading‑based deployments. It uses layer‑wise exit supervision, self‑speculative decoding over a Top‑K local vocabulary, and dynamic hidden‑state management to enable reliable intermediate‑layer predictions and memory‑aware execution. Experiments on Llama2‑7B and Llama3‑8B show that FlexEE achieves significant speedups—up to 1.27×/3.16× and 1.25×/2.83× respectively—while maintaining minimal accuracy loss.

By Qihu Xie, Ziwei Li, Yi Kang