arXiv AI By Siva Rajesh Kasa, Yasong Dai, Sumit Negi, Hongdong Li

Fast-dLLM++: Fr\'{e}chet Profile Decoding for Faster Diffusion LLM Inference

Read the original on arXiv AI →

arXiv:2606. 02955v1 Announce Type: cross Abstract: Diffusion large language models promise parallel token generation, yet inference remains bottlenecked by deciding which masked tokens can be safely committed together.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Reliable Parallel Decoding in Masked Diffusion Language Models

The paper introduces Reliable Parallel Decoding (RPD) for masked diffusion language models, addressing the unreliability of committing multiple predictions from a single forward pass. Diagnostics reveal that confidence alone is insufficient, as confident end‑sequence predictions can preempt necessary upstream computations, and downstream predictions degrade with upstream uncertainty. RPD selects candidates based on layer‑wise stability and final confidence, committing them under an entropy budget while deferring uncertain predictions, achieving superior throughput and competitive accuracy on LLaDA and Dream benchmarks.

By Zhenghao He, Bohan Liu, Guangzhi Xiong, Aidong Zhang
arXiv Computation and Language
Sep 23

PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models

The paper introduces PACE-dLLM, an acceleration method for diffusion language models (dLLMs) that uses the model’s own per‑step confidence to estimate a ‘confidence cliff’ and determine the optimal look‑ahead horizon for block decoding. By fitting this cliff in closed form at each step, PACE-dLLM sets the horizon to its saturation point and applies an independent confidence threshold for token commitment, thereby avoiding the trade‑offs inherent in fixed‑size block decoding. Experiments on reasoning and code benchmarks show that PACE-dLLM achieves the best average accuracy on open‑source dLLM backbones while delivering significant wall‑clock speedups—up to 5.23× on LLaDA and 3.06× on Dream—improving the quality‑throughput Pareto frontier.

By Xiaocheng Lu, Shuhan Guo, Ziyue Ma, Jie Zhang, Jian Liu, Jingcai Guo, Haoxuan Che, Song Guo