Ripple-Pivot Search (RPS) is a training‑free decoding method for Diffusion Large Language Models that identifies mid‑entropy pivot positions to reduce uncertainty across remaining masked tokens. By proactively committing these pivots and evaluating token assignments via lookahead, RPS enables more tokens to be unmasked in parallel, speeding up decoding. Experiments on three dLLMs and four reasoning/code‑generation benchmarks show 4–10× wall‑clock speedup over standard decoding, up to 18× with KV caching, while maintaining or improving generation quality.
By Yushi Ye, Xu Chen, Haoyun Jiang, Jinsong Lan, Haihong Tang, Xiangtao Li, Mingming Gong, Ivor Tsang, Yanfeng Wang, Jiangchao Yao
arXiv:2604. 18995v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive generation by enabling parallel token prediction.
By Zhenbang Du, Kejing Xia, Xinrui Zhong, Yonggan Fu, Nicolai Oswald, Binfei Ji, Brucek Khailany, Pavlo Molchanov, Yingyan Lin
arXiv:2607. 15655v1 Announce Type: cross Abstract: Masked diffusion language models (DLMs) enable parallel text generation by iteratively refining masked tokens, offering a promising alternative to autoregressive decoding.
By Yingqian Cui, Wei Deng, Lantao Mei, Hang Li, Charu C. Aggarwal, Hui Liu, Yue Xing
arXiv:2606. 04236v1 Announce Type: cross Abstract: Discrete diffusion language models can generate text efficiently by updating multiple masked positions in parallel, but this parallelism introduces a quality-latency trade-off.
By Giries Abu Ayoub, Mario Barbara, Llu\'is Pastor-P\'erez, Tanja Bien, Aneesh Barthakur, Alaa Maalouf, Loay Mualem
arXiv:2609.16450v1 Announce Type: cross
Abstract: Diffusion large language models (dLLMs) offer a promising parallel decoding paradigm as an alternative to autoregressive generation through iterative...
By Lixuan Wei, Wei Zhou, Jianwen Wu, Yipeng Shen, Meiling Wang, Haoran You
arXiv:2607. 20467v1 Announce Type: new Abstract: While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds.
By Yanhua Jiao, Tianyi Wu, Xiaoxi Sun, Yulin Li, HuiLing Zhen, Libo Qin, Baotian Hu, Zhuotao Tian, Min Zhang
The paper introduces Reliable Parallel Decoding (RPD) for masked diffusion language models, addressing the unreliability of committing multiple predictions from a single forward pass. Diagnostics reveal that confidence alone is insufficient, as confident end‑sequence predictions can preempt necessary upstream computations, and downstream predictions degrade with upstream uncertainty. RPD selects candidates based on layer‑wise stability and final confidence, committing them under an entropy budget while deferring uncertain predictions, achieving superior throughput and competitive accuracy on LLaDA and Dream benchmarks.
By Zhenghao He, Bohan Liu, Guangzhi Xiong, Aidong Zhang
arXiv:2601. 17917v3 Announce Type: replace Abstract: Diffusion Large Language Models (dLLMs) offer a compelling paradigm for natural language generation, leveraging parallel decoding and bidirectional attention to achieve superior global coherence compared to autoregressive models.
By Zhongyu Xiao, Zhiwei Hao, Jianyuan Guo, Yong Luo, Jia Liu, Jie Xu, Han Hu
arXiv:2606. 02544v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decoding.
By Junxia Cui, Haotian Ye, Runchu Tian, Hongcan Guo, Jinya Jiang, Haoru Li, Chaojie Ren, Yiming Huang, Kaijie Zhu, Zhongkai Yu, Kun Zhou, Jingbo Shang
arXiv:2607. 14107v1 Announce Type: cross Abstract: The inference efficiency of diffusion large language models (dLLMs) is constrained by two challenges: bidirectional attention precludes efficient KV-cache reuse, while increasing decoding parallelism with static confidence thresholds can compromise generation quality.
By Mingyu Lee, Akshat Ramachandran, Souvik Kundu, Tushar Krishna
The paper introduces Dependency-Aware Revocable Decoding (DARD), a training‑free framework for diffusion large language models that separates tokens into masked, candidate, and unmasked states. DARD verifies candidate tokens using a selective context that excludes less reliable tokens and adaptively regulates their influence on subsequent decoding. Experiments on 12 textual and multimodal benchmarks across three open‑source dLLMs show that DARD improves the speed‑quality Pareto frontier, achieving a 2.71× speedup and a 4.35‑point CIDEr gain over Saber on Flickr30K.
By Wooje Park, Insu Lee, Minyoung Noh, Jaeyun Jang, Sungmin Lee, Kyuhong Shim, Byonghyo Shim
arXiv:2606. 15805v1 Announce Type: new Abstract: Discrete diffusion language models enable parallel token generation, offering a pathway to low-latency decoding.
By Tamim Zoabi, Ameen Ali, Liran Ringel, Lior Wolf