arXiv AI By Zhenghao He, Bohan Liu, Guangzhi Xiong, Aidong Zhang

Reliable Parallel Decoding in Masked Diffusion Language Models

Read the original on arXiv AI →

The paper introduces Reliable Parallel Decoding (RPD) for masked diffusion language models, addressing the unreliability of committing multiple predictions from a single forward pass. Diagnostics reveal that confidence alone is insufficient, as confident end‑sequence predictions can preempt necessary upstream computations, and downstream predictions degrade with upstream uncertainty. RPD selects candidates based on layer‑wise stability and final confidence, committing them under an entropy budget while deferring uncertain predictions, achieving superior throughput and competitive accuracy on LLaDA and Dream benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 21

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

Ripple-Pivot Search (RPS) is a training‑free decoding method for Diffusion Large Language Models that identifies mid‑entropy pivot positions to reduce uncertainty across remaining masked tokens. By proactively committing these pivots and evaluating token assignments via lookahead, RPS enables more tokens to be unmasked in parallel, speeding up decoding. Experiments on three dLLMs and four reasoning/code‑generation benchmarks show 4–10× wall‑clock speedup over standard decoding, up to 18× with KV caching, while maintaining or improving generation quality.

By Yushi Ye, Xu Chen, Haoyun Jiang, Jinsong Lan, Haihong Tang, Xiangtao Li, Mingming Gong, Ivor Tsang, Yanfeng Wang, Jiangchao Yao
Hugging Face Trending Papers
Aug 12

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster inference through parallel decoding. Existing parallel decoding schedulers typically commit positions only after they meet a per-position criterion, overlooking how early commitments may benefit subsequent decoding.