A Survey on Diffusion Language Models
arXiv:2508. 10875v3 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm.
The paper introduces Hierarchical Continuous Diffusion Language Models (H-CDLMs), a framework that jointly diffuses multiple token modalities—individual tokens and coarser clusters of token embeddings—to enhance continuous diffusion language models. Applied to the CoBit architecture, the resulting H-CoBit achieves significant empirical gains, improving MAUVE scores and achieving lower generative perplexity on LM1B and OWT, while also outperforming prior continuous diffusion models on GSM8K. The approach generalizes to other continuous generative paradigms, as shown by consistent improvements when applied to the flow matching model FLM.
arXiv:2508. 10875v3 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm.
Hierarchical Continuous Diffusion Language Models (HC‑DLM) combine discrete token generation with a continuous latent trajectory in a single denoising process, addressing the independence bottleneck of parallel decoding in discrete diffusion models. The training objective is derived from a variational bound on token likelihood, and the latent state is the sole persistent generative element, with tokens read out and fed back at each step. Experiments on Sudoku, Countdown, and LM1B show HC‑DLM outperforming both discrete and continuous diffusion baselines in puzzle accuracy and generative perplexity.
arXiv:2601. 22954v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel.
The paper introduces Sigma, a large-scale continuous diffusion language model (3B/8B parameters) that uses steerable, low-dimensional ODE/SDE latent trajectories to address the non-smoothness of discrete diffusion models. Sigma is trained blockwise via likelihood optimization, jointly denoises Gaussian-corrupted token embeddings, and learns an optimal embedding geometry, leveraging pre-trained autoregressive weights for faster training. During inference, classifier-free guidance and score temperature are identified as essential for high-fidelity reasoning and coding, and Sigma matches or exceeds discrete models on benchmarks such as GSM8K, Minerva, HumanEval, MBPP, MATH-500, and AIME, while also revealing unique structural benefits like embedding-space steering and graceful degradation for low NFEs.
arXiv:2609.15070v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, independently predicting multiple tokens at each step. This...
arXiv:2609.40362v1 Announce Type: new Abstract: We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and q...
arXiv:2606. 19475v1 Announce Type: new Abstract: Large Language Models (LLMs) have revolutionized language modeling through autoregressive generation, enabling strong performance across a wide range of tasks.
The paper introduces a training‑free structured suffix modeling technique to accelerate Diffusion Language Models (DLMs). It partitions the suffix into local, middle, and tail regions, retaining varying numbers of tokens per region and incorporating previous decoding results into current token representations. Experiments on three DLMs show significant speedups—up to 72.81× in long‑sequence inference—while often improving performance, and the method is compatible with existing acceleration strategies.
arXiv:2608.23551v1 Announce Type: cross Abstract: Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existi...
PlaidQ is a 0.7B continuous diffusion language model designed for code generation. By distilling its iterative refinement trajectory into only a few denoising steps—or even a single step—PlaidQ achieves competitive performance with discrete diffusion models while dramatically reducing inference time. The study demonstrates that continuous diffusion can be effectively compressed, enabling efficient and accurate code generation with minimal computational overhead.
arXiv:2609. 04010v1 Announce Type: new Abstract: Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation.
The paper introduces Window-Diffusion, a method that accelerates diffusion language model inference by pruning and caching tokens within a sliding window. It categorizes undecoded tokens into active, buffer, and far-field groups, computing only the first two while discarding the rest. Experiments on LLaDA and Dream demonstrate up to 99× speedup with minimal loss in generation quality.