The paper proposes a new covariance model for Denoising Diffusion Probabilistic Models (DDPMs) that captures non‑diagonal correlations and the power‑law frequency spectrum of natural images. Using a Kronecker‑factored DCT (K‑DCT) decomposition, the authors reduce computational complexity from quadratic to log‑linear, enabling efficient sampling with few steps. Experiments on CIFAR‑10, Celeb‑A, ImageNet, and LSUN demonstrate improved FID and likelihoods over previous state‑of‑the‑art samplers.
By Rui Xia, Ayan Das, Artem Artemev, Andi Zhang, Guillaume Hennequin, Alberto Bernacchia
arXiv:2609.00955v1 Announce Type: new
Abstract: Diffusion models achieve strong image generation quality but incur high iterative denoising costs. Analog compute-in-memory (CIM) can accelerate matrix...
By Yuannuo Feng, Yizhe Chen, Wenshuai Yao, Yuxin Xie, Ngai Wong, Wenyong Zhou, Wang Kang
arXiv:2609.39648v1 Announce Type: cross
Abstract: Diffusion models are typically viewed as stochastic processes that transform noise into data. We take a complementary perspective: a diffusion model...
By Cristina L\'opez Amado, Marco Fumero, Francesco Locatello
arXiv:2605.21907v2 Announce Type: replace
Abstract: Test-Time Scaling (TTS) paradigm offers a promising perspective for enhancing the generation performance of diffusion models. However, current solu...
By Gang Dai, Yining Huang, Yiming Xia, Guohao Chen, Shuaicheng Niu
arXiv:2602. 02908v2 Announce Type: replace-cross Abstract: Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same noise seed.
By Binxu Wang, Jacob Zavatone-Veth, Cengiz Pehlevan
The paper argues that the concrete random noise used in diffusion models is not merely a passive perturbation but a learnable input that can be exploited by the model. By analyzing how clean data and realized noise jointly form the noisy input, the authors show that the model can learn regularities in the data or in the noise structure, and that these two routes can interact. Experiments on MNIST and CIFAR‑10 using pseudorandom streams demonstrate that structured‑noise training can reduce prediction loss, but this advantage disappears when test noise is replaced with IID noise, indicating that the learned dependence is tied to the specific noise structure.
By Shengzhi Deng, Chenqi Ye, Yanze Guo
Spatially Adaptive Noise Injection (SANI) is a new diffusion sampling framework that adjusts the amount of noise added at each pixel during reverse diffusion. Unlike traditional samplers that apply a uniform noise variance across the image, SANI uses a probabilistic gating mechanism to inject noise only where the denoiser is uncertain, such as edges and textures, while preserving smooth regions. Experiments show that SANI consistently improves Fréchet Inception Distance over vanilla DDPM and DDIM samplers across various timesteps, and remains competitive with variance‑learning baselines.
By Frantzeska Lavda, Maciej Falkiewicz, Van Khoa Nguyen, Alexandros Kalousis
arXiv:2606. 27696v1 Announce Type: cross Abstract: In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion models.
By Jiequan Cui, Beier Zhu, Qingshan Xu, Xiaojuan Qi, Bei Yu, Hanwang Zhang
arXiv:2602. 17706v2 Announce Type: replace Abstract: Diffusion models learn data distributions indirectly through denoising, making the difficulty of generative modeling closely tied to the dependency structure of data.
By Rongyao Cai, Yuxi Wan, Kexin Zhang, Ming Jin, Zhiqiang Ge, Qingsong Wen, Yong Liu
arXiv:2606. 02661v1 Announce Type: cross Abstract: Accurate precipitation nowcasting is vital for disaster mitigation, but deep learning methods face a key trade-off: regression models produce over-smoothed, spectrally decaying predictions that blur convective details and violate turbulence power laws; diffusion models generate realistic yet unanchored hallucinations lacking physical grounding.
By Yunlong Zhou, Chen Zhao, Danyang Peng, Fanfan Ji, Xiao-Tong Yuan
arXiv:2603. 14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility?
By Chujun Tang, Lei Zhong, Fangqiang Ding
Existing diffusion-based methods have recently made significant progress in image dehazing. However, they typically neglect the physics of haze formation and reconstruct clean images from pure Gaussian noise, thereby limiting their restoration potential.