arXiv Machine Learning By Yair Schiff, Omer Belhasin, Roy Uziel, Matan Rusanovsky, Ran Zilberstein, Marianne Arriola, Gilad Turok, Guanghan Wang, Volodymyr Kuleshov, Michael Elad

Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models

Read the original on arXiv Machine Learning →

The paper introduces Clock Diffusion, a framework for semi‑autoregressive continuous diffusion language models that incorporates position‑dependent noise schedules, efficient training, and sampling algorithms. It presents two generation modes—block and sliding window—and defines ClockDLMs, a family of Gaussian models that achieve state‑of‑the‑art diffusion likelihoods on OpenWebText and outperform continuous baselines on GSM8K while matching or exceeding discrete diffusion models. The authors also propose Cache Grab, a set of efficient samplers that adapt accelerated inference techniques from discrete diffusion to further improve model quality and efficiency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 4

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

arXiv:2609. 04010v1 Announce Type: new Abstract: Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation.

By Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter, Chaitanya Dwivedi, Varad Pimpalkhute, Yash Akhauri, Alexander Moreno, Mikhail Yurochkin, Zhenting Wang, Mostafa Elhoushi, Nolan Dey, Shane Bergsma, Joel Hestness, John Thickstun, Eric Xing, Zhengzhong Liu
arXiv AI
Sep 18

Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference

Zarya is a hybrid language model that jointly trains an autoregressive objective and a masked-diffusion objective within a single architecture. It structures training data into variable-size slots and uses a curriculum that gradually increases slot granularity, allowing a smooth transition from fine-grained AR learning to coarse-grained diffusion learning. At inference, Zarya offers two decoding modes—MDM sampling with first-hitting denoising and slotted speculative decoding that interleaves diffusion-based selection with autoregressive infilling—while fully decoupling training and inference regimes and supporting extensive configurability.

By Leonid Sinev, Ilya Koziev, Vladislav Leshchuk
arXiv AI
Jul 17

Token Time Continuous Diffusion for Language Modeling

arXiv:2607. 14106v1 Announce Type: cross Abstract: In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates in continuous space, deterministically mapping Gaussian noise to a final token canvas with no further sampling, and crucially (b) incorporates a new notion of per-token times, with some tokens proceeding from noise to token at a faster rate than others.

By Parikshit Bansal, Sujay Sanghavi
arXiv Machine Learning
Sep 11

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

The paper revisits the continuous diffusion language model Plaid and introduces RePlaid, aligning its architecture with modern discrete diffusion models. RePlaid achieves a compute gap of only 20× compared to autoregressive models, surpasses Duo with fewer parameters, and outperforms MDLM in over‑trained settings. On OpenWebText, RePlaid sets a new state‑of‑the‑art continuous diffusion perplexity of 22.1 and demonstrates superior generation quality, while theoretical analysis links likelihood‑based training to linear cross‑entropy over time and structured embedding geometries.

By Zhihan Yang, Wei Guo, Shuibai Zhang, Subham Sekhar Sahoo, Yongxin Chen, Arash Vahdat, Morteza Mardani, John Thickstun