arXiv Machine Learning
Jun 2

Consistent Diffusion Language Models

arXiv:2605. 00161v2 Announce Type: replace Abstract: Diffusion language models (DLMs) are an attractive alternative to autoregressive models because they promise sublinear-time, parallel generation, yet practical gains remain elusive as high-quality samples still demand hundreds of refinement steps.

By Hasan Amin, Yuan Gao, Yaser Souri, Subhojit Som, Ming Yin, Rajiv Khanna, Xia Song
arXiv AI
Sep 10

Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks

The paper examines how Denoising Diffusion Probabilistic Models (DDPMs) perform on globally constrained discrete tasks such as Sudoku and N-queens. It shows that standard diffusion sampling, which keeps updates close to the noisy state, often preserves early mistakes, whereas sampling directly from the model’s clean predictions dramatically improves validity (e.g., Sudoku from 31% to 95%). The authors further introduce self‑correction training, exposing the model to its own predictions to reduce inference errors, which enhances the performance of standard samplers across tasks.

By Mariia Drozdova, St\'ephane Liem Nguyen, Fran\c{c}ois Fleuret
arXiv Machine Learning
Aug 20

Foundations of Diffusion Models in General State Spaces: A Self-Contained Introduction

The article "Foundations of Diffusion Models in General State Spaces: A Self-Contained Introduction" presents a unified primer on diffusion models that applies to both continuous Euclidean data and discrete categorical structures. It develops discrete-time forward noising via Markov kernels and learned reverse dynamics, and connects these to continuous-time limits such as stochastic differential equations in ρ^d and continuous-time Markov chains on finite alphabets, deriving the corresponding Fokker–Planck and master equations. The work also shows how different forward corruption choices—Gaussian processes for continuous spaces and structured categorical transition kernels for discrete spaces—affect reverse dynamics and the evidence lower bound used in training, offering a layered exposition for newcomers, practitioners, and experts alike.

By Vincent Pauline, Tobias H\"oppe, Kirill Neklyudov, Alexander Tong, Stefan Bauer, Andrea Dittadi
arXiv AI
Jul 28

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

arXiv:2607. 24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling.

By Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv AI
4d ago

Less Uniform Discrete Diffusion is More Powerful and Scalable

The paper introduces Less Uniform Diffusion (LUDI), a framework that improves uniform diffusion language models by using a less uniform loss and per-token time embeddings to guide reverse transitions and enable confidence-based few-step sampling. Experiments demonstrate that LUDI provides cleaner supervision, enhances few-step generation, and scales to a 7B model (LUDI-7B) that achieves a 3-token-per-step speedup over autoregressive decoding while matching masked diffusion baselines. The work suggests that UDLMs still have untapped potential for complex generation tasks.

By Kaibo Wang, Ding Ding, Fangyu Ding, Zijin Feng, Han Shi, Haili Bai, Jiacheng Sun, Yang Xiang