arXiv Computation and Language

Conditional Total Correlation and the Serial Depth of Adaptive Parallel Sampling

The paper introduces a new framework for adaptive parallel sampling of discrete vectors, where a deterministic policy reveals coordinates round‑by‑round based on previously observed values and samples the remaining coordinates from their exact conditional marginals. The authors prove an exact identity linking the forward Kullback‑Leibler divergence of any policy to the expected conditional total correlation accumulated during the sampling process, establishing conditional total correlation as the precise information cost of within‑round parallelism. Using this identity, they derive zero‑error schedules for finite‑order Markov chains, characterize the serial depth of Bernoulli walks, and demonstrate separations between different reveal orders, permutation strategies, and string structures, thereby revealing how conditional dependence governs parallelizability. whyItMatters":"The results provide a principled, information‑theoretic measure of parallel sampling efficiency that can guide the design of decoding rules for masked diffusion models and other generative systems."

arXiv Machine Learning
Jun 24

ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs

arXiv:2510. 04767v2 Announce Type: replace Abstract: While most autoregressive LLMs are constrained to one-by-one decoding, diffusion LLMs (dLLMs) have attracted growing interest for their potential to dramatically accelerate inference through parallel decoding.

By Wonjun Kang, Kevin Galim, Seunghyuk Oh, Minjae Lee, Yuchen Zeng, Shuibai Zhang, Coleman Hooper, Yuezhou Hu, Hyung Il Koo, Nam Ik Cho, Kangwook Lee
arXiv Machine Learning
Sep 18

Parallelism, critical windows, and separations among diffusion language models

The paper compares the parallelism capabilities of three diffusion large language model paradigms—masked, uniform, and Gaussian diffusion. It proves that uniform and Gaussian diffusion can sample with a number of forward passes scaling with the dual total correlation of the distribution, potentially much less than the context length, whereas masked diffusion may require more passes. The study establishes a provable separation in parallelism, showing that masked diffusion’s critical windows are asymptotically narrower than those of the other two approaches.

By Sitan Chen, Liye Wang
arXiv AI
Jun 10

Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

arXiv:2606. 10829v1 Announce Type: cross Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled.

By Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro
arXiv Machine Learning
Sep 22

Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories

The paper studies how to allocate a fixed computational budget across the denoising steps of diffusion models to improve sample quality at deployment. It shows that the expected benefit of evaluating multiple candidates at a step can be decomposed into a step‑specific sensitivity and a universal sample‑size factor, and that the optimal allocation follows a water‑filling structure. Experiments demonstrate that this allocation achieves the same quality as a uniform strategy while reducing function evaluations by 20–50%.

By Yuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng