arXiv Machine Learning

Adaptive Order Policies for Masked Diffusion

arXiv:2606. 00295v1 Announce Type: new Abstract: Masked diffusion models have seen great success in capturing data distributions over discrete sequences in domains such as text and proteins.

arXiv Computation and Language
Sep 23

Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

Informed Masking (IM) is a new technique for aligning Diffusion Large Language Models (dLLMs) with Reinforcement Learning (RL). It identifies a systematic upstream/downstream token structure in dLLM rollouts and shows that masking downstream tokens creates better subproblems for likelihood estimation. When integrated into three state‑of‑the‑art dLLM RL methods on LLaDA‑8B‑Instruct, IM yields up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks while improving training stability.

By Xiaoyi Yu, Enver Sangineto, Pei Fu, Fiorenzo Parascandolo, Wenhui Tan, Ruikang Zhang, Rita Cucchiara, Ruihua Song, Jian Luan
arXiv Machine Learning
Jun 3

Learning Unmasking Policies for Diffusion Language Models

arXiv:2512. 09106v4 Announce Type: replace Abstract: Diffusion (Large) Language Models (dLLMs) now match the downstream performance of their autoregressive counterparts on many tasks, while holding the promise of being more efficient during inference.

By Metod Jazbec, Theo X. Olausson, Louis B\'ethune, Pierre Ablin, Michael Kirchhof, Jo\~ao Monteiro, Victor Turrisi, Jason Ramapuram, Marco Cuturi
arXiv Machine Learning
Aug 31

Learned Relay Representations for Forward-Thinking Discrete Diffusion Models

The paper introduces Learned Relay Representations (Relay), a technique for Masked Diffusion Models (MDMs) that preserves and forwards internal latent information across denoising steps via a differentiable per-token channel trained with truncated backpropagation through time. Relay enables MDMs to be forward‑thinking, avoiding costly recomputation of internal representations and improving efficiency. Applied to Fast‑dLLM v2, Relay outperforms standard supervised fine‑tuning on coding tasks and reduces inference latency by up to 32%, demonstrating a clear performance‑latency advantage for diffusion language models.

By Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel, Neil Band, Avishek Joey Bose, Tim G. J. Rudner, Andrew McCallum
arXiv Machine Learning
Jun 9

Insertion Based Sequence Generation with Learnable Order Dynamics

arXiv:2602. 18695v2 Announce Type: replace Abstract: Existing insertion-based masked diffusion models that generate sequences by interleaving token insertion with unmasking use fixed schedules that are not dependent on the data.

By Dhruvesh Patel, Benjamin Rozonoyer, Gaurav Pandey, Tahira Naseem, Ram\'on Fernandez Astudillo, Andrew McCallum