The paper studies how to allocate a fixed computational budget across the denoising steps of diffusion models to improve sample quality at deployment. It shows that the expected benefit of evaluating multiple candidates at a step can be decomposed into a step‑specific sensitivity and a universal sample‑size factor, and that the optimal allocation follows a water‑filling structure. Experiments demonstrate that this allocation achieves the same quality as a uniform strategy while reducing function evaluations by 20–50%.
By Yuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng
arXiv:2605.08144v2 Announce Type: replace-cross
Abstract: Training a diffusion model involves two sources of randomness for each data sample: the timestep and the Gaussian noise realization. The time...
By Haokai Zhao, Da Xing, Hanqun Cao, Tinson Xu, Xinyu Xiang, Yanchao Li, Xiangru Tang, Hongbin Lin, Zehong Wang, Kuan Pang, Peng Xia, Molei Tao, Li Erran Li, Aditya Joshi, Jure Leskovec, Fang Wu
arXiv:2601. 22450v2 Announce Type: replace-cross Abstract: Masked Diffusion Language Models have recently emerged as a powerful generative paradigm, yet their generalization properties remain understudied compared to their auto-regressive counterparts.
By Jianhao Huang, Baharan Mirzasoleiman
arXiv:2606. 13796v1 Announce Type: cross Abstract: Recursive training of generative models on their own outputs can lead to model collapse, a compounding drift away from the true data distribution.
By Na\"il B. Khelifa, Richard E. Turner, Ramji Venkataramanan
arXiv:2605.12597v3 Announce Type: replace-cross
Abstract: Computational sampling has been central to the sciences since the mid-20th century. While machine-learning-based approaches have recently ena...
By Luca Maria Del Bono, Giulio Biroli, Patrick Charbonneau, Marylou Gabri\'e
The paper argues that the concrete random noise used in diffusion models is not merely a passive perturbation but a learnable input that can be exploited by the model. By analyzing how clean data and realized noise jointly form the noisy input, the authors show that the model can learn regularities in the data or in the noise structure, and that these two routes can interact. Experiments on MNIST and CIFAR‑10 using pseudorandom streams demonstrate that structured‑noise training can reduce prediction loss, but this advantage disappears when test noise is replaced with IID noise, indicating that the learned dependence is tied to the specific noise structure.
By Shengzhi Deng, Chenqi Ye, Yanze Guo