arXiv Machine Learning

Adaptive Block Diffusion: Resolving Training-Inference Mismatch in Diffusion Language Models

arXiv:2606. 29275v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) are typically trained under fixed context structures, restricting denoising to predetermined token subsets.

arXiv AI
Jun 15

Residual Context Diffusion Language Models

arXiv:2601. 22954v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel.

By Yuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi, Coleman Hooper, Jintao Zhang, Aditya Tomar, Michael W. Mahoney, Sewon Min, Mehrdad Farajtabar, Kurt Keutzer, Amir Gholami, Chenfeng Xu
arXiv Machine Learning
Jun 25

Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding

arXiv:2601. 17917v3 Announce Type: replace Abstract: Diffusion Large Language Models (dLLMs) offer a compelling paradigm for natural language generation, leveraging parallel decoding and bidirectional attention to achieve superior global coherence compared to autoregressive models.

By Zhongyu Xiao, Zhiwei Hao, Jianyuan Guo, Yong Luo, Jia Liu, Jie Xu, Han Hu
arXiv AI
Sep 3

Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models

The paper introduces PILL, a new infilling technique for diffusion language models that eliminates the need for a preset initial length and reduces inference overhead. PILL uses probing-based length-free decoding, cutting down on extra forward passes and speeding up generation. Experiments across five diffusion models and eight benchmarks show PILL outperforms the strongest baseline with higher pass rates and BLEU-2 scores while running 1.82× faster.

By Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong
arXiv Machine Learning
Jun 3

Learning Unmasking Policies for Diffusion Language Models

arXiv:2512. 09106v4 Announce Type: replace Abstract: Diffusion (Large) Language Models (dLLMs) now match the downstream performance of their autoregressive counterparts on many tasks, while holding the promise of being more efficient during inference.

By Metod Jazbec, Theo X. Olausson, Louis B\'ethune, Pierre Ablin, Michael Kirchhof, Jo\~ao Monteiro, Victor Turrisi, Jason Ramapuram, Marco Cuturi
arXiv AI
Sep 18

Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference

Zarya is a hybrid language model that jointly trains an autoregressive objective and a masked-diffusion objective within a single architecture. It structures training data into variable-size slots and uses a curriculum that gradually increases slot granularity, allowing a smooth transition from fine-grained AR learning to coarse-grained diffusion learning. At inference, Zarya offers two decoding modes—MDM sampling with first-hitting denoising and slotted speculative decoding that interleaves diffusion-based selection with autoregressive infilling—while fully decoupling training and inference regimes and supporting extensive configurability.

By Leonid Sinev, Ilya Koziev, Vladislav Leshchuk
arXiv Computation and Language
Sep 23

Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

Informed Masking (IM) is a new technique for aligning Diffusion Large Language Models (dLLMs) with Reinforcement Learning (RL). It identifies a systematic upstream/downstream token structure in dLLM rollouts and shows that masking downstream tokens creates better subproblems for likelihood estimation. When integrated into three state‑of‑the‑art dLLM RL methods on LLaDA‑8B‑Instruct, IM yields up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks while improving training stability.

By Xiaoyi Yu, Enver Sangineto, Pei Fu, Fiorenzo Parascandolo, Wenhui Tan, Ruikang Zhang, Rita Cucchiara, Ruihua Song, Jian Luan
arXiv AI
Jun 3

$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction

arXiv:2604. 18995v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive generation by enabling parallel token prediction.

By Zhenbang Du, Kejing Xia, Xinrui Zhong, Yonggan Fu, Nicolai Oswald, Binfei Ji, Brucek Khailany, Pavlo Molchanov, Yingyan Lin
arXiv Machine Learning
Sep 16

Window-Diffusion: Accelerating Diffusion Language Model Inference with Windowed Token Pruning and Caching

The paper introduces Window-Diffusion, a method that accelerates diffusion language model inference by pruning and caching tokens within a sliding window. It categorizes undecoded tokens into active, buffer, and far-field groups, computing only the first two while discarding the rest. Experiments on LLaDA and Dream demonstrate up to 99× speedup with minimal loss in generation quality.

By Fengrui Zuo, Zhiwei Ke, Yiming Liu, Wenqi Lou, Chao Wang, Xuehai Zhou