arXiv AI

Simple Self-Conditioning Adaptation for Masked Diffusion Models

arXiv:2604. 26985v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) generate discrete sequences by iterative denoising under an absorbing masking process.

arXiv AI
Jul 28

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

arXiv:2607. 24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling.

By Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv AI
Jun 10

MMD Guidance: Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy Guidance

arXiv:2601. 08379v2 Announce Type: replace-cross Abstract: Pre-trained diffusion models have emerged as powerful generative priors for both unconditional and conditional sample generation, yet their outputs often deviate from the characteristics of user-specific target data.

By Matina Mahdizadeh Sani, Nima Jamali, Mohammad Jalali, Farzan Farnia
arXiv AI
Sep 10

Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks

The paper examines how Denoising Diffusion Probabilistic Models (DDPMs) perform on globally constrained discrete tasks such as Sudoku and N-queens. It shows that standard diffusion sampling, which keeps updates close to the noisy state, often preserves early mistakes, whereas sampling directly from the model’s clean predictions dramatically improves validity (e.g., Sudoku from 31% to 95%). The authors further introduce self‑correction training, exposing the model to its own predictions to reduce inference errors, which enhances the performance of standard samplers across tasks.

By Mariia Drozdova, St\'ephane Liem Nguyen, Fran\c{c}ois Fleuret
arXiv Computer Vision
Aug 26

On-Policy Self-Distillation in Diffusion Models

The paper introduces DiffusionOPSD, an on‑policy self‑distillation framework that transforms image‑level reinforcement learning rewards into explicit targets for intermediate denoising predictions in diffusion models. By generating trajectories with a frozen behavior policy and constructing bounded positive and negative targets around query states, the method trains a policy to fit these targets before updating the behavior policy via an exponential moving average. Experiments on SD 3.5‑M and Z‑Image‑Turbo show that DiffusionOPSD achieves the best held‑out scores in 19 of 20 reward‑matched settings, outperforms the strongest competitor by up to 44 % and cuts GPU‑hour usage by 40–63 % compared to DiffusionNFT.

By Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
arXiv Computation and Language
Sep 23

Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

Informed Masking (IM) is a new technique for aligning Diffusion Large Language Models (dLLMs) with Reinforcement Learning (RL). It identifies a systematic upstream/downstream token structure in dLLM rollouts and shows that masking downstream tokens creates better subproblems for likelihood estimation. When integrated into three state‑of‑the‑art dLLM RL methods on LLaDA‑8B‑Instruct, IM yields up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks while improving training stability.

By Xiaoyi Yu, Enver Sangineto, Pei Fu, Fiorenzo Parascandolo, Wenhui Tan, Ruikang Zhang, Rita Cucchiara, Ruihua Song, Jian Luan
arXiv Machine Learning
Aug 31

Learned Relay Representations for Forward-Thinking Discrete Diffusion Models

The paper introduces Learned Relay Representations (Relay), a technique for Masked Diffusion Models (MDMs) that preserves and forwards internal latent information across denoising steps via a differentiable per-token channel trained with truncated backpropagation through time. Relay enables MDMs to be forward‑thinking, avoiding costly recomputation of internal representations and improving efficiency. Applied to Fast‑dLLM v2, Relay outperforms standard supervised fine‑tuning on coding tasks and reduces inference latency by up to 32%, demonstrating a clear performance‑latency advantage for diffusion language models.

By Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel, Neil Band, Avishek Joey Bose, Tim G. J. Rudner, Andrew McCallum