arXiv Computation and Language By Chuling Wen, Weijie Liang, Jian Lu

Conditional Total Correlation and the Serial Depth of Adaptive Parallel Sampling

Read the original on arXiv Computation and Language →

The paper introduces a new framework for adaptive parallel sampling of discrete vectors, where a deterministic policy reveals coordinates round‑by‑round based on previously observed values and samples the remaining coordinates from their exact conditional marginals. The authors prove an exact identity linking the forward Kullback‑Leibler divergence of any policy to the expected conditional total correlation accumulated during the sampling process, establishing conditional total correlation as the precise information cost of within‑round parallelism. Using this identity, they derive zero‑error schedules for finite‑order Markov chains, characterize the serial depth of Bernoulli walks, and demonstrate separations between different reveal orders, permutation strategies, and string structures, thereby revealing how conditional dependence governs parallelizability. whyItMatters":"The results provide a principled, information‑theoretic measure of parallel sampling efficiency that can guide the design of decoding rules for masked diffusion models and other generative systems."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jun 24

ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs

arXiv:2510. 04767v2 Announce Type: replace Abstract: While most autoregressive LLMs are constrained to one-by-one decoding, diffusion LLMs (dLLMs) have attracted growing interest for their potential to dramatically accelerate inference through parallel decoding.

By Wonjun Kang, Kevin Galim, Seunghyuk Oh, Minjae Lee, Yuchen Zeng, Shuibai Zhang, Coleman Hooper, Yuezhou Hu, Hyung Il Koo, Nam Ik Cho, Kangwook Lee
arXiv Machine Learning
Sep 18

Parallelism, critical windows, and separations among diffusion language models

The paper compares the parallelism capabilities of three diffusion large language model paradigms—masked, uniform, and Gaussian diffusion. It proves that uniform and Gaussian diffusion can sample with a number of forward passes scaling with the dual total correlation of the distribution, potentially much less than the context length, whereas masked diffusion may require more passes. The study establishes a provable separation in parallelism, showing that masked diffusion’s critical windows are asymptotically narrower than those of the other two approaches.

By Sitan Chen, Liye Wang