arXiv Machine Learning

Neither Parallel Nor Sequential: How DiffusionGemma Actually Commits Tokens

arXiv:2606. 14620v1 Announce Type: new Abstract: Open diffusion language models are marketed as parallel, non-autoregressive decoders, yet the order in which a shipped checkpoint actually commits its tokens is almost never measured.

arXiv AI
Aug 25

Answer First, Reason Later: When Commitment Order Costs Accuracy in Diffusion Language Models

The paper studies how the order in which tokens are committed in masked diffusion language models affects accuracy. It finds that when the final answer is committed before the preceding reasoning (an answer‑first trajectory), accuracy can suffer compared to unrestricted decoding, especially on tasks like GSM8K and MATH‑500. Experiments with controlled token positions show that delaying the answer token can improve performance, indicating that commitment order influences the context and output allocation of the model.

By Jewon Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Hwiyeong Lee, Taesup Kim
arXiv AI
Jun 19

How Transparent is DiffusionGemma?

arXiv:2606. 20560v1 Announce Type: cross Abstract: LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors.

By Joshua Engels, Callum McDougall, Bilal Chughtai, Janos Kramar, Senthoran Rajamanoharan, Cindy Wu, Arthur Conmy, Asic Q Chen, Jean Tarbouriech, Min Ma, Brendan O'Donoghue, Jo\~ao Gabriel Lopes de Oliveira, Rohin Shah, Neel Nanda
Hugging Face Trending Papers
Jun 10

Teaching Diffusion to Speculate Left-to-Right

Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inherently sequential token generation. Speculative decoding addresses this bottleneck by employing a lightweight draft model to propose multiple future tokens that are subsequently verified in parallel by a larger target model.

arXiv Machine Learning
Sep 18

Parallelism, critical windows, and separations among diffusion language models

The paper compares the parallelism capabilities of three diffusion large language model paradigms—masked, uniform, and Gaussian diffusion. It proves that uniform and Gaussian diffusion can sample with a number of forward passes scaling with the dual total correlation of the distribution, potentially much less than the context length, whereas masked diffusion may require more passes. The study establishes a provable separation in parallelism, showing that masked diffusion’s critical windows are asymptotically narrower than those of the other two approaches.

By Sitan Chen, Liye Wang