arXiv:2509. 21474v4 Announce Type: replace Abstract: While diffusion language models (DLMs) have achieved competitive performance in text generation, improving their reasoning ability with reinforcement learning remains an active research area.
By Guanghan Wang, Gilad Turok, Yair Schiff, Marianne Arriola, Volodymyr Kuleshov
Informed Masking (IM) is a new technique for aligning Diffusion Large Language Models (dLLMs) with Reinforcement Learning (RL). It identifies a systematic upstream/downstream token structure in dLLM rollouts and shows that masking downstream tokens creates better subproblems for likelihood estimation. When integrated into three state‑of‑the‑art dLLM RL methods on LLaDA‑8B‑Instruct, IM yields up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks while improving training stability.
By Xiaoyi Yu, Enver Sangineto, Pei Fu, Fiorenzo Parascandolo, Wenhui Tan, Ruikang Zhang, Rita Cucchiara, Ruihua Song, Jian Luan
arXiv:2510. 04019v3 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) represent a promising alternative to autoregressive LLMs; however, the lack of effective post-training techniques, including reinforcement learning (RL), remains a key challenge for dLLMs, especially for downstream applications.
By Anthony Zhan
The paper introduces Less Uniform Diffusion (LUDI), a framework that improves uniform diffusion language models by using a less uniform loss and per-token time embeddings to guide reverse transitions and enable confidence-based few-step sampling. Experiments demonstrate that LUDI provides cleaner supervision, enhances few-step generation, and scales to a 7B model (LUDI-7B) that achieves a 3-token-per-step speedup over autoregressive decoding while matching masked diffusion baselines. The work suggests that UDLMs still have untapped potential for complex generation tasks.
By Kaibo Wang, Ding Ding, Fangyu Ding, Zijin Feng, Han Shi, Haili Bai, Jiacheng Sun, Yang Xiang
arXiv:2609.16372v1 Announce Type: new
Abstract: Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention. Extending reasoning across...
By Albert Ge, Chandan Singh, Yufan Zhuang, Xiaodong Liu, Jianfeng Gao, Frederic Sala
arXiv:2606. 02544v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decoding.
By Junxia Cui, Haotian Ye, Runchu Tian, Hongcan Guo, Jinya Jiang, Haoru Li, Chaojie Ren, Yiming Huang, Kaijie Zhu, Zhongkai Yu, Kun Zhou, Jingbo Shang