arXiv:2607. 26285v1 Announce Type: cross Abstract: Two central challenges in diffusion-based sampling are the theoretical one of understanding their remarkable effectiveness even in high-dimensional settings, and the practical one of designing algorithms with certified performance guarantees.
By Martin J. Wainwright
arXiv:2608.28949v1 Announce Type: cross
Abstract: We study a class of product-reference diffusion algorithms for sampling from a discrete distribution. We show that their sampling performance can be...
By Martin J. Wainwright
arXiv:2602. 15008v2 Announce Type: replace Abstract: Diffusion models over discrete spaces have recently shown striking empirical success, yet their theoretical foundations remain incomplete.
By Daniil Dmitriev, Zhihan Huang, Yuting Wei
The paper introduces a new framework for adaptive parallel sampling of discrete vectors, where a deterministic policy reveals coordinates round‑by‑round based on previously observed values and samples the remaining coordinates from their exact conditional marginals. The authors prove an exact identity linking the forward Kullback‑Leibler divergence of any policy to the expected conditional total correlation accumulated during the sampling process, establishing conditional total correlation as the precise information cost of within‑round parallelism. Using this identity, they derive zero‑error schedules for finite‑order Markov chains, characterize the serial depth of Bernoulli walks, and demonstrate separations between different reveal orders, permutation strategies, and string structures, thereby revealing how conditional dependence governs parallelizability.
whyItMatters":"The results provide a principled, information‑theoretic measure of parallel sampling efficiency that can guide the design of decoding rules for masked diffusion models and other generative systems."
By Chuling Wen, Weijie Liang, Jian Lu
The paper studies how to choose sampling schedules for tau‑leaping in masked discrete diffusion models. By deriving an exact integral representation of the factorization error ε_fact in terms of a dependence density ρ, the authors develop estimators and recursive equations that identify the unique optimal schedule under a monotonicity condition. In the large‑scale limit, they provide explicit characterizations of the optimal smooth schedule and show that while optimizing smooth schedules can improve constants, it does not change the N/K scaling unless the dependence density degenerates, in which case asymptotic improvements are possible.
By Cecilia Secchi, Giacomo Zanella
The paper proves that for discrete diffusion models using uniform or remasking forward processes, an adaptive sampler based on a leave‑one‑out denoiser can achieve sampling error proportional to the score‑estimation error plus a small tolerance. The required number of discretization steps scales with the dual total correlation of the target distribution, not directly with the ambient dimension. This result shows that sampling complexity is governed by the intrinsic dependence structure of the distribution, and the authors provide an information‑theoretic analysis linking discretization error to mutual information between coordinates.
By Daniil Dmitriev, Zhihan Huang, Yuting Wei
arXiv:2512. 24152v2 Announce Type: replace-cross Abstract: Sampling based on score diffusions has led to striking empirical results, and has attracted considerable attention from various research communities.
By M. J. Wainwright
The paper studies how to allocate a fixed computational budget across the denoising steps of diffusion models to improve sample quality at deployment. It shows that the expected benefit of evaluating multiple candidates at a step can be decomposed into a step‑specific sensitivity and a universal sample‑size factor, and that the optimal allocation follows a water‑filling structure. Experiments demonstrate that this allocation achieves the same quality as a uniform strategy while reducing function evaluations by 20–50%.
By Yuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng
arXiv:2606. 23867v1 Announce Type: new Abstract: The exact computation of the Normalized Maximum Likelihood (NML) codelength for regular non-smooth estimators (e.
By Trenton Lau, Gary P. T. Choi
The paper presents a non‑asymptotic analysis of Markov chain Monte Carlo (MCMC) algorithms that learn and apply a preconditioner based on either the target covariance or the expected Hessian of the target potential. It compares the finite‑time computational costs of these preconditioned schemes with unpreconditioned counterparts, providing guarantees for algorithms such as the Unadjusted Langevin Algorithm (ULA) and the proximal sampler. The analysis relies on a contraction assumption in the Wasserstein‑2 distance to formalize approximate independence and bridge modern MCMC theory with classical effective sample size heuristics.
By Max Hird, Florian Maire, Jeffrey Negrea
arXiv:2605. 17232v5 Announce Type: replace Abstract: Discrete diffusion has become a leading framework for generative modeling in various applications including language, vision, and biology.
By Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Markos A. Katsoulakis
arXiv:2607. 00773v1 Announce Type: new Abstract: Discrete diffusion models are widely used for learning and generating discrete distributions.
By Yu Yao, Huanjian Zhou, Andi Han, Wei Huang, Masashi Sugiyama