arXiv Machine Learning By Manvi Jha, Zach Zhang, Zhichao Xu, Linbo Liu, Sai Muralidhar Jayanthi, Vinayak Arannil

APEX: Speculate smarter, not deeper

Read the original on arXiv Machine Learning →

The paper introduces APEX, a learned controller that improves speculative decoding for large language models by dynamically selecting the best speculation strategy and adjusting draft depth during generation. APEX uses request‑level routing among EAGLE‑3, n‑gram, and draft‑model speculation, and block‑level depth adaptation based on causal decoding signals and verifier feedback. Integrated into vLLM and tested with Qwen3‑8B, APEX achieves up to 5.24× speedup over autoregressive decoding while reducing wasted tokens by 41% compared to fixed‑depth n‑gram speculation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 4

Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding

The paper introduces AdaptiveSpec, a training‑free speculative decoding method that simultaneously adapts the per‑step verification rule and the draft‑tree shape using signals generated during decoding. It replaces the fixed token‑match rule with a margin‑based threshold and adjusts tree depth, width, and node count based on draft confidence and recent acceptance history, allowing the total draft count to vary. Experiments on SGLang show up to 56% throughput gains over EAGLE‑3 while maintaining 93% of lossless task accuracy on GSM8K, MATH‑500, and HumanEval across three models.

By Oszk\'ar Urb\'an, Young D. Kwon, Stylianos I. Venieris, Cecilia Mascolo
arXiv Computation and Language
1d ago

DLoop: Looped Speculative Decoding

The paper introduces DLoop, a looped speculative decoding technique that adaptively performs multiple drafting stages before verification, allowing a draft model to continue generating tokens while confident. By verifying all accumulated draft tokens together and training the draft model to handle its own hidden states for unverified tokens, DLoop reduces the number of target‑model forward passes needed. Experiments across several speculative decoding methods show wall‑clock speedups of 5–41 % without sacrificing lossless decoding.

By Geonmo Gu, Byeongho Heo, HeeJae Jun, Yoohoon Kang, Sangmin Lee, Sangdoo Yun, Dongyoon Han
arXiv Computation and Language
Sep 23

TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models

The paper introduces TSS, a target-side sparsification framework that selectively skips layers in a target verifier during speculative decoding for domain-specific large language models. By exploring multi-layer skip configurations with an acceptance- and metric-aware breadth search, TSS reduces verification cost, increases draft acceptance, and can even improve downstream task performance without retraining. Experiments on Spec-Bench demonstrate consistent gains across domains and model scales, notably boosting translation throughput by 1.68× and improving BLEU scores significantly.

By Haibo Hu, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue