arXiv AI

BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration

arXiv Machine Learning
Sep 21

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.

By Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T
Hugging Face Trending Papers
Jul 27

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.

arXiv AI
Jul 7

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

arXiv:2607. 05147v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification.

By Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang
arXiv Computation and Language
Sep 23

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Flash-dLLM is a training‑free inference acceleration framework that improves the speed and memory efficiency of Diffusion Large Language Models (dLLMs). It tackles GPU memory I/O bottlenecks by introducing an I/O‑aware fused KV‑cache kernel and then employs a draft‑and‑verify decoding strategy that uses the dLLM itself as both drafter and verifier. Experiments on mathematical reasoning and code‑generation tasks show Flash‑dLLM outperforms existing acceleration methods, achieving up to 11.0× speedups over the Elastic‑Cache baseline.

By Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
arXiv Computation and Language
Sep 23

TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models

The paper introduces TSS, a target-side sparsification framework that selectively skips layers in a target verifier during speculative decoding for domain-specific large language models. By exploring multi-layer skip configurations with an acceptance- and metric-aware breadth search, TSS reduces verification cost, increases draft acceptance, and can even improve downstream task performance without retraining. Experiments on Spec-Bench demonstrate consistent gains across domains and model scales, notably boosting translation throughput by 1.68× and improving BLEU scores significantly.

By Haibo Hu, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue