arXiv AI

Fast-dLLM++: Fr\'{e}chet Profile Decoding for Faster Diffusion LLM Inference

arXiv:2606. 02955v1 Announce Type: cross Abstract: Diffusion large language models promise parallel token generation, yet inference remains bottlenecked by deciding which masked tokens can be safely committed together.

arXiv AI
3d ago

Reliable Parallel Decoding in Masked Diffusion Language Models

The paper introduces Reliable Parallel Decoding (RPD) for masked diffusion language models, addressing the unreliability of committing multiple predictions from a single forward pass. Diagnostics reveal that confidence alone is insufficient, as confident end‑sequence predictions can preempt necessary upstream computations, and downstream predictions degrade with upstream uncertainty. RPD selects candidates based on layer‑wise stability and final confidence, committing them under an entropy budget while deferring uncertain predictions, achieving superior throughput and competitive accuracy on LLaDA and Dream benchmarks.

By Zhenghao He, Bohan Liu, Guangzhi Xiong, Aidong Zhang
arXiv Computation and Language
Sep 23

PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models

The paper introduces PACE-dLLM, an acceleration method for diffusion language models (dLLMs) that uses the model’s own per‑step confidence to estimate a ‘confidence cliff’ and determine the optimal look‑ahead horizon for block decoding. By fitting this cliff in closed form at each step, PACE-dLLM sets the horizon to its saturation point and applies an independent confidence threshold for token commitment, thereby avoiding the trade‑offs inherent in fixed‑size block decoding. Experiments on reasoning and code benchmarks show that PACE-dLLM achieves the best average accuracy on open‑source dLLM backbones while delivering significant wall‑clock speedups—up to 5.23× on LLaDA and 3.06× on Dream—improving the quality‑throughput Pareto frontier.

By Xiaocheng Lu, Shuhan Guo, Ziyue Ma, Jie Zhang, Jian Liu, Jingcai Guo, Haoxuan Che, Song Guo
arXiv Computation and Language
Sep 21

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

Ripple-Pivot Search (RPS) is a training‑free decoding method for Diffusion Large Language Models that identifies mid‑entropy pivot positions to reduce uncertainty across remaining masked tokens. By proactively committing these pivots and evaluating token assignments via lookahead, RPS enables more tokens to be unmasked in parallel, speeding up decoding. Experiments on three dLLMs and four reasoning/code‑generation benchmarks show 4–10× wall‑clock speedup over standard decoding, up to 18× with KV caching, while maintaining or improving generation quality.

By Yushi Ye, Xu Chen, Haoyun Jiang, Jinsong Lan, Haihong Tang, Xiangtao Li, Mingming Gong, Ivor Tsang, Yanfeng Wang, Jiangchao Yao
arXiv AI
Jun 10

Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

arXiv:2606. 10829v1 Announce Type: cross Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled.

By Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro
Hugging Face Trending Papers
Aug 12

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster inference through parallel decoding. Existing parallel decoding schedulers typically commit positions only after they meet a per-position criterion, overlooking how early commitments may benefit subsequent decoding.

arXiv AI
Jun 2

SimSD: Simple Speculative Decoding in Diffusion Language Models

arXiv:2606. 02544v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decoding.

By Junxia Cui, Haotian Ye, Runchu Tian, Hongcan Guo, Jinya Jiang, Haoru Li, Chaojie Ren, Yiming Huang, Kaijie Zhu, Zhongkai Yu, Kun Zhou, Jingbo Shang