arXiv AI By Dake Bu, Wei Huang, Andi Han, Hau-San Wong, Qingfu Zhang, Taiji Suzuki, Atsushi Nitanda

DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models

Read the original on arXiv AI →

arXiv:2604. 24357v2 Announce Type: replace-cross Abstract: Diffusion language models generate without a fixed left-to-right order, leaving token ordering as a central algorithmic choice.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 25

DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Discrete Diffusion Models

The paper introduces DPRM, a plug‑in token‑ordering module for discrete diffusion models that uses a Doob h‑transform to convert terminal rewards into per‑position process rewards. By estimating these rewards from generation progress, confidence, and optional state, DPRM reorders tokens without altering the underlying model or sampler. Experiments on nine open‑source hosts show significant gains in reasoning, numeric VQA, visual‑codebook ordering, and preference‑conditioned generation, with improvements ranging from 8.97 to 53.3 points across diverse tasks.

By Dake Bu, Wei Huang, Andi Han, Si Wu, Hau-San Wong, Qingfu Zhang, Taiji Suzuki, Atsushi Nitanda
arXiv Computation and Language
Sep 23

Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

Informed Masking (IM) is a new technique for aligning Diffusion Large Language Models (dLLMs) with Reinforcement Learning (RL). It identifies a systematic upstream/downstream token structure in dLLM rollouts and shows that masking downstream tokens creates better subproblems for likelihood estimation. When integrated into three state‑of‑the‑art dLLM RL methods on LLaDA‑8B‑Instruct, IM yields up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks while improving training stability.

By Xiaoyi Yu, Enver Sangineto, Pei Fu, Fiorenzo Parascandolo, Wenhui Tan, Ruikang Zhang, Rita Cucchiara, Ruihua Song, Jian Luan
Hugging Face Trending Papers
Jun 10

Teaching Diffusion to Speculate Left-to-Right

Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inherently sequential token generation. Speculative decoding addresses this bottleneck by employing a lightweight draft model to propose multiple future tokens that are subsequently verified in parallel by a larger target model.