Speculative decoding leverages idle CPU resources to accelerate token generation without altering model outputs. In vLLM benchmarks, DFlash achieved a 3.92× increase in autoregressive throughput using Qwen3.5‑9B on an Intel Xeon 6 at a concurrency of 1. The article details the origins of this speedup, discusses acceptance metrics, and outlines factors that influence when speculation is beneficial.
By Ehssan Khan
arXiv:2602. 07223v2 Announce Type: replace Abstract: Long-context large language model (LLM) inference has become the norm for today's AI applications.
By Yikang Yue, Yuqi Xue, Jian Huang
The study benchmarks six quantum error‑correction decoders on the Willow processor, the first device operating below the surface‑code threshold, using a hierarchy of increasingly realistic noise models. By evaluating real hardware data across multiple code distances, bases, and round counts, the authors find that rank agreement with hardware emerges only when each operation type is assigned its own error rate. They also independently test NVIDIA’s Ising pre‑decoder, showing it offers no accuracy‑latency advantage over other decoders in most evaluations, and release the full evaluation pipeline and data for future comparisons.
By Shay J. Manor, Leila S. Erhili, Yassine Jebbouri
arXiv:2606. 08382v1 Announce Type: cross Abstract: Low-rank projection has emerged as a promising approach for compressing the KV cache by exploiting hidden-dimension redundancy.
By Priyansh Bhatnagar, Ashkan Moradifirouzabadi, Se-Hyun Yang, SeungJae Lee, Jungwook Choi, Mingu Kang
The paper introduces SPECTRA, a runtime‑reconfigurable tiled architecture designed to accelerate speculative decoding for large language models on edge devices. SPECTRA adapts its compute engine within each tile between systolic GEMM execution and vector‑lane GEMV execution, while dynamically adjusting tile count, kernel partitioning, and communication patterns across tiles. Experiments on a 20‑tile FPGA prototype demonstrate up to a 2.09× speedup from tile‑level reconfiguration and an additional 1.25× improvement from system‑level adaptability compared to fixed designs.
By Gabriele Tombesi, William Baisi, Je Yang, Elisavet Lydia Alvanaki, Kevin Lee, Michael Lippe, Biruk Seyoum, Luca P. Carloni
arXiv:2607. 04244v1 Announce Type: new Abstract: This report describes our approach to the Efficient Qwen Competition, where the goal is to enable low-latency serving of Qwen3.
By Jaeyeon Kim, Jewon Lee, Bo-Kyeong Kim
arXiv:2607. 22389v1 Announce Type: cross Abstract: With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck.
By Chao Fang, Jun Yin, Man Shi, Marian Verhelst
arXiv:2512. 22420v5 Announce Type: replace-cross Abstract: Speculative decoding (SD) accelerates LLM inference by verifying draft tokens in parallel.
By Rui Li, Zhaoning Zhang, Libo Zhang, Huaimin Wang, Xiang Fu, Zhiquan Lai
SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.
By Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T