Towards Data Science
Aug 24

Speculative Decoding on CPUs: Nearly 4x Faster Token Generation with DFlash

Speculative decoding leverages idle CPU resources to accelerate token generation without altering model outputs. In vLLM benchmarks, DFlash achieved a 3.92× increase in autoregressive throughput using Qwen3.5‑9B on an Intel Xeon 6 at a concurrency of 1. The article details the origins of this speedup, discusses acceptance metrics, and outlines factors that influence when speculation is beneficial.

By Ehssan Khan
arXiv Machine Learning
Sep 7

A Sim-to-Real Study of Surface-Code Decoder Benchmarking

The study benchmarks six quantum error‑correction decoders on the Willow processor, the first device operating below the surface‑code threshold, using a hierarchy of increasingly realistic noise models. By evaluating real hardware data across multiple code distances, bases, and round counts, the authors find that rank agreement with hardware emerges only when each operation type is assigned its own error rate. They also independently test NVIDIA’s Ising pre‑decoder, showing it offers no accuracy‑latency advantage over other decoders in most evaluations, and release the full evaluation pipeline and data for future comparisons.

By Shay J. Manor, Leila S. Erhili, Yassine Jebbouri