The paper introduces DPRM, a plug‑in token‑ordering module for discrete diffusion models that uses a Doob h‑transform to convert terminal rewards into per‑position process rewards. By estimating these rewards from generation progress, confidence, and optional state, DPRM reorders tokens without altering the underlying model or sampler. Experiments on nine open‑source hosts show significant gains in reasoning, numeric VQA, visual‑codebook ordering, and preference‑conditioned generation, with improvements ranging from 8.97 to 53.3 points across diverse tasks.
By Dake Bu, Wei Huang, Andi Han, Si Wu, Hau-San Wong, Qingfu Zhang, Taiji Suzuki, Atsushi Nitanda
Informed Masking (IM) is a new technique for aligning Diffusion Large Language Models (dLLMs) with Reinforcement Learning (RL). It identifies a systematic upstream/downstream token structure in dLLM rollouts and shows that masking downstream tokens creates better subproblems for likelihood estimation. When integrated into three state‑of‑the‑art dLLM RL methods on LLaDA‑8B‑Instruct, IM yields up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks while improving training stability.
By Xiaoyi Yu, Enver Sangineto, Pei Fu, Fiorenzo Parascandolo, Wenhui Tan, Ruikang Zhang, Rita Cucchiara, Ruihua Song, Jian Luan
Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inherently sequential token generation. Speculative decoding addresses this bottleneck by employing a lightweight draft model to propose multiple future tokens that are subsequently verified in parallel by a larger target model.
arXiv:2606. 11552v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inherently sequential token generation.
By Lexington Whalen, Yuki Ito, Ryo Sakamoto
arXiv:2609.38806v1 Announce Type: cross
Abstract: Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on prob...
By Woosang Jeon, Jaeyeon Kim, Sham Kakade, Yilun Du, Amrit Singh Bedi, Arun Kumar Chithanar, Chul Lee, Taehyeong Kim, Sitan Chen
arXiv:2601. 15165v4 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary orders.
By Zanlin Ni, Shenzhi Wang, Yang Yue, Tianyu Yu, Weilin Zhao, Yeguo Hua, Tianyi Chen, Jun Song, Cheng Yu, Bo Zheng, Gao Huang