arXiv AI By Chia-Ming Lee, Shao-Kai Liu, Ming-Ching Chang, Xin Li, Yu-Lun Liu, Chih-Chung Hsu

Commit Locally, Exit Globally: Coordinating Adaptive Sampling and Early Exit in Diffusion Language Models

Read the original on arXiv AI →

arXiv:2607. 28166v2 Announce Type: replace-cross Abstract: Diffusion language models expose a provisional prediction at every denoising step, and on many tasks the candidate answer inside it stabilizes before the step schedule is exhausted.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models

The paper introduces PACE-dLLM, an acceleration method for diffusion language models (dLLMs) that uses the model’s own per‑step confidence to estimate a ‘confidence cliff’ and determine the optimal look‑ahead horizon for block decoding. By fitting this cliff in closed form at each step, PACE-dLLM sets the horizon to its saturation point and applies an independent confidence threshold for token commitment, thereby avoiding the trade‑offs inherent in fixed‑size block decoding. Experiments on reasoning and code benchmarks show that PACE-dLLM achieves the best average accuracy on open‑source dLLM backbones while delivering significant wall‑clock speedups—up to 5.23× on LLaDA and 3.06× on Dream—improving the quality‑throughput Pareto frontier.

By Xiaocheng Lu, Shuhan Guo, Ziyue Ma, Jie Zhang, Jian Liu, Jingcai Guo, Haoxuan Che, Song Guo
arXiv AI
Sep 2

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.

By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim