Improved Large Language Diffusion Models
arXiv:2606. 25331v1 Announce Type: cross Abstract: Modern large language models are predominantly trained with autoregressive factorization and causal attention.
arXiv:2606. 25331v1 Announce Type: cross Abstract: Modern large language models are predominantly trained with autoregressive factorization and causal attention.
arXiv:2511. 21338v2 Announce Type: replace Abstract: Masked Diffusion Language Models (MDLMs) have recently emerged as a promising alternative to Autoregressive Language Models (ARLMs), leveraging a denoising objective that, in principle, should enable more uniform context utilisation.
arXiv:2606. 04236v1 Announce Type: cross Abstract: Discrete diffusion language models can generate text efficiently by updating multiple masked positions in parallel, but this parallelism introduces a quality-latency trade-off.
arXiv:2607. 15200v1 Announce Type: cross Abstract: Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation.
Diffusion large language models (dLLMs) offer an efficient alternative to autoregressive models through parallel decoding, yet existing post-training methods largely rely on random masking strategies that overlook intrinsic token dependencies. In this work, we present an empirical analysis of attention in dLLMs and show that tokens attending more strongly to unmasked context exhibit greater generation stability and play a critical role in reasoning.
arXiv:2604. 18738v3 Announce Type: replace Abstract: Diffusion language models (dLLMs) generate text through iterative denoising, filling multiple masked positions at each step.
arXiv:2601. 22954v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel.
arXiv:2609.37974v1 Announce Type: cross Abstract: Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The m...
arXiv:2609.36202v1 Announce Type: cross Abstract: Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time....
arXiv:2604. 18995v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive generation by enabling parallel token prediction.
arXiv:2606. 10829v1 Announce Type: cross Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled.
MDM-Prime-v2 is a masked diffusion language model that improves upon the original MDM-Prime by addressing two key limitations: the high cross‑entropy loss caused by common Byte‑Pair‑Encoding tokenizers and the lack of guidance for choosing token granularity. The new design incorporates Binary Encoding and Index Shuffling, and the authors provide an analysis linking token granularity and sub‑token entropy to training objectives and downstream performance. When scaled to 1.1 B parameters, MDM-Prime‑v2 achieves higher average zero‑shot accuracy on eight commonsense reasoning benchmarks than comparable models such as GPT‑Neo, OPT, Pythia, Bloom, SMDM, and TinyLLaMA.