arXiv Machine Learning
Sep 7

Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One

PlaidQ is a 0.7B continuous diffusion language model designed for code generation. By distilling its iterative refinement trajectory into only a few denoising steps—or even a single step—PlaidQ achieves competitive performance with discrete diffusion models while dramatically reducing inference time. The study demonstrates that continuous diffusion can be effectively compressed, enabling efficient and accurate code generation with minimal computational overhead.

By Fred Zhangzhi Peng, Kaiwen Zheng, Anru R. Zhang
arXiv Machine Learning
Sep 4

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

arXiv:2609. 04010v1 Announce Type: new Abstract: Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation.

By Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, Zhengzhong Liu
Hugging Face Trending Papers
Jun 10

Teaching Diffusion to Speculate Left-to-Right

Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inherently sequential token generation. Speculative decoding addresses this bottleneck by employing a lightweight draft model to propose multiple future tokens that are subsequently verified in parallel by a larger target model.