Pixel‑Space Diffusion via Observation Operators introduces a new framework for pixel‑space diffusion models that addresses a scale‑time mismatch in existing methods. By replacing fixed full‑image supervision with a time‑indexed observation trajectory that progresses from coarse structures to the full image, the model aligns supervision with the natural recovery order of image details. The approach employs Gaussian‑Lanczos operators and a GL‑CoDA decoder to refine features progressively, resulting in faster convergence and higher generation quality, achieving an FID of 1.52 on ImageNet‑256.
By Shaojie Guo, Lichen Ma, Haoyang Tong, Yu He, Zipeng Guo, Xiaoan Liu, Feng Yan, Yu Guo, Fei Wang, Junshi Huang, Yan Wang
arXiv:2608.29322v1 Announce Type: new
Abstract: Recent video diffusion models have achieved remarkable generation quality, but high-fidelity results still largely depend on closed-source systems or c...
By Hangzhou He, Lunhao Duan, Shanshan Zhao, Kaiwen Li, Qing-Guo Chen, Weihua Luo, Yanye Lu
The paper introduces a Posterior‑Dynamics Framework that leverages pretrained diffusion models as multiscale priors for linear imaging inverse problems such as deblurring, super‑resolution, and inpainting. By constructing a surrogate likelihood centered on the clean image and incorporating diffusion uncertainty, the authors derive continuous posterior dynamics and a tunable Langevin component for adaptive exploration. They prove theoretical guarantees (endpoint consistency, finite‑horizon tracking, weak accuracy) and present the PD‑IMEX sampler, which achieves high‑quality reconstructions with only 100 score evaluations and controllable fidelity‑diversity trade‑offs.
By Zhaoqiang Liu, Tongyao Pang, Ruibing Wang, Yang Zheng
arXiv:2603. 14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility?
By Chujun Tang, Lei Zhong, Fangqiang Ding
arXiv:2607. 09753v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success across diverse domains, with performance closely related to the denoising backbones that parameterize the score function.
By Haksoo Lim, Myeongjin Lee, Wonjoon Chang, Jaesik Choi
The paper introduces CM-RED, a fast MRI reconstruction method that combines a pretrained consistency model with the regularization by denoising framework. By integrating controlled noise injection into accelerated proximal gradient updates, CM-RED achieves high‑quality reconstructions on fastMRI knee and brain datasets with only four network function evaluations. It consistently outperforms existing diffusion‑ and consistency‑based approaches in quantitative metrics, visual fidelity, and robustness to hyperparameter changes.
By Merve G\"ulle, Junno Yun, Ya\c{s}ar Utku Al\c{c}alar, Mehmet Ak\c{c}akaya
arXiv:2608. 15144v1 Announce Type: cross Abstract: Posterior sampling with a pretrained diffusion prior is governed by a conditional score whose intermediate likelihood component is generally intractable.
By Zhaoqiang Liu, Tongyao Pang, Ruibing Wang, Yang Zheng
AlignMorph is a tuning‑free diffusion framework for image morphing that separates geometric alignment from generative denoising. It uses Global Semantic Transport—entropic optimal transport and reliability‑aware latent warping—to achieve diffusion‑compatible semantic alignment, and Coordinate‑Aligned Generation—symmetric bi‑phase attention handoff—to preserve spatial coordinates during denoising. The method eliminates ghosting and delivers superior structural coherence and temporal smoothness on morphing benchmarks without any per‑pair optimization.
By Wuyi Liu, Xu Han, Yuren Chen, Yige Mao, Zishuo Peng, Xianzhi Li
arXiv:2609.00798v1 Announce Type: new
Abstract: Pixel-space diffusion has recently emerged as a promising direction for high-fidelity image generation by modeling images directly in the original pixe...
By Weiyi You, Jinhua Zhang, Xingyu Zhou, Wei Long, Junyu Lou, Shuhang Gu
Pixel-space diffusion has recently emerged as a promising direction for high-fidelity image generation by modeling images directly in the original pixel domain. However, pixel-space diffusion is compu...
arXiv:2604.24136v3 Announce Type: replace
Abstract: Pretrained diffusion models have revolutionized real-world image super-resolution (Real-ISR), but their iterative sampling is computationally prohi...
By Shyang-En Weng, Yi-Cheng Liao, Yu-Syuan Xu, Chia-Hung Yuan, Wei-Chen Chiu, Ching-Chun Huang
arXiv:2607. 05319v1 Announce Type: cross Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures.
By Rajat Rasal, Avinash Kori, Tian Xia, Ben Glocker