Hugging Face Blog

Accelerate StarCoder with 🤗 Optimum Intel on Xeon: Q8/Q4 and Speculative Decoding

Towards Data Science
Aug 24

Speculative Decoding on CPUs: Nearly 4x Faster Token Generation with DFlash

Speculative decoding leverages idle CPU resources to accelerate token generation without altering model outputs. In vLLM benchmarks, DFlash achieved a 3.92× increase in autoregressive throughput using Qwen3.5‑9B on an Intel Xeon 6 at a concurrency of 1. The article details the origins of this speedup, discusses acceptance metrics, and outlines factors that influence when speculation is beneficial.

By Ehssan Khan
arXiv Machine Learning
Sep 7

A Sim-to-Real Study of Surface-Code Decoder Benchmarking

The study benchmarks six quantum error‑correction decoders on the Willow processor, the first device operating below the surface‑code threshold, using a hierarchy of increasingly realistic noise models. By evaluating real hardware data across multiple code distances, bases, and round counts, the authors find that rank agreement with hardware emerges only when each operation type is assigned its own error rate. They also independently test NVIDIA’s Ising pre‑decoder, showing it offers no accuracy‑latency advantage over other decoders in most evaluations, and release the full evaluation pipeline and data for future comparisons.

By Shay J. Manor, Leila S. Erhili, Yassine Jebbouri
arXiv AI
Sep 23

SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled Architecture

The paper introduces SPECTRA, a runtime‑reconfigurable tiled architecture designed to accelerate speculative decoding for large language models on edge devices. SPECTRA adapts its compute engine within each tile between systolic GEMM execution and vector‑lane GEMV execution, while dynamically adjusting tile count, kernel partitioning, and communication patterns across tiles. Experiments on a 20‑tile FPGA prototype demonstrate up to a 2.09× speedup from tile‑level reconfiguration and an additional 1.25× improvement from system‑level adaptability compared to fixed designs.

By Gabriele Tombesi, William Baisi, Je Yang, Elisavet Lydia Alvanaki, Kevin Lee, Michael Lippe, Biruk Seyoum, Luca P. Carloni
arXiv Machine Learning
Sep 21

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.

By Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T