arXiv AI

The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity

arXiv:2608. 13520v1 Announce Type: cross Abstract: We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \emph{unmasking growth complexity} ({\textsf{UGC}\xspace}).

arXiv Computation and Language
Aug 27

Conditional Total Correlation and the Serial Depth of Adaptive Parallel Sampling

The paper introduces a new framework for adaptive parallel sampling of discrete vectors, where a deterministic policy reveals coordinates round‑by‑round based on previously observed values and samples the remaining coordinates from their exact conditional marginals. The authors prove an exact identity linking the forward Kullback‑Leibler divergence of any policy to the expected conditional total correlation accumulated during the sampling process, establishing conditional total correlation as the precise information cost of within‑round parallelism. Using this identity, they derive zero‑error schedules for finite‑order Markov chains, characterize the serial depth of Bernoulli walks, and demonstrate separations between different reveal orders, permutation strategies, and string structures, thereby revealing how conditional dependence governs parallelizability. whyItMatters":"The results provide a principled, information‑theoretic measure of parallel sampling efficiency that can guide the design of decoding rules for masked diffusion models and other generative systems."

By Chuling Wen, Weijie Liang, Jian Lu
arXiv Machine Learning
Sep 21

Schedule optimization for tau-leaping in masked discrete diffusion

The paper studies how to choose sampling schedules for tau‑leaping in masked discrete diffusion models. By deriving an exact integral representation of the factorization error ε_fact in terms of a dependence density ρ, the authors develop estimators and recursive equations that identify the unique optimal schedule under a monotonicity condition. In the large‑scale limit, they provide explicit characterizations of the optimal smooth schedule and show that while optimizing smooth schedules can improve constants, it does not change the N/K scaling unless the dependence density degenerates, in which case asymptotic improvements are possible.

By Cecilia Secchi, Giacomo Zanella
arXiv Statistics ML
Aug 25

Provably adaptive sampling with uniform and remasking discrete diffusion models

The paper proves that for discrete diffusion models using uniform or remasking forward processes, an adaptive sampler based on a leave‑one‑out denoiser can achieve sampling error proportional to the score‑estimation error plus a small tolerance. The required number of discretization steps scales with the dual total correlation of the target distribution, not directly with the ambient dimension. This result shows that sampling complexity is governed by the intrinsic dependence structure of the distribution, and the authors provide an information‑theoretic analysis linking discretization error to mutual information between coordinates.

By Daniil Dmitriev, Zhihan Huang, Yuting Wei
arXiv Machine Learning
Sep 22

Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories

The paper studies how to allocate a fixed computational budget across the denoising steps of diffusion models to improve sample quality at deployment. It shows that the expected benefit of evaluating multiple candidates at a step can be decomposed into a step‑specific sensitivity and a universal sample‑size factor, and that the optimal allocation follows a water‑filling structure. Experiments demonstrate that this allocation achieves the same quality as a uniform strategy while reducing function evaluations by 20–50%.

By Yuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng
arXiv Statistics ML
Aug 26

A Non-asymptotic Analysis for Learning and Applying a Preconditioner in MCMC

The paper presents a non‑asymptotic analysis of Markov chain Monte Carlo (MCMC) algorithms that learn and apply a preconditioner based on either the target covariance or the expected Hessian of the target potential. It compares the finite‑time computational costs of these preconditioned schemes with unpreconditioned counterparts, providing guarantees for algorithms such as the Unadjusted Langevin Algorithm (ULA) and the proximal sampler. The analysis relies on a contraction assumption in the Wasserstein‑2 distance to formalize approximate independence and bridge modern MCMC theory with classical effective sample size heuristics.

By Max Hird, Florian Maire, Jeffrey Negrea