arXiv:2608.20515v1 Announce Type: new
Abstract: Generative video compression can recover rich visual details at low bitrates, but simultaneously achieving high temporal consistency and low inference...
By Wenzhuo Ma, Zhenzhong Chen
arXiv:2608.29243v1 Announce Type: new
Abstract: Existing diffusion-based enhancement methods provide strong generative capability for low-light image enhancement (LLIE), yet they either rely on paire...
By Wenjie Cai, Yuezhe Yang, Jianyang Xia, Xingbo Dong, Zhe Jin
The paper introduces a perceptually regularized diffusion framework for image super‑resolution, adding perceptual‑loss based regularization to the standard diffusion training objective. This approach incorporates prior knowledge to improve training convergence and encourages the recovery of meaningful image features. Experiments on benchmark datasets show enhanced perceptual quality while maintaining competitive distortion metrics.
By Chuxiangbo Wang, Pavithra Venkatachalapathy, Ying Liang, Min Wang, Jing Qin, Yifei Lou, Weihong Guo
arXiv:2607. 06856v1 Announce Type: cross Abstract: Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics.
By Michael King, Aravindh Mahendran, Matthew Koichi Grimes, Fedor Kitashov, Adham Elarabawy, Pedro Velez, Maks Ovsjanikov, Viorica P\u{a}tr\u{a}ucean
The paper introduces an adaptive step schedule controller for text‑to‑image diffusion models, allowing the number of denoising steps to vary based on the complexity of the input prompt. By mixing step schedules of different sizes and monitoring error discrepancies at each timestep, the method switches schedules to maintain image quality while reducing inference time. Experiments on COCO and DiffusionDB demonstrate that this approach achieves faster generation without sacrificing visual fidelity.
By Kuluhan Binici, Cihan Acar, Shivam Aggarwal, Siying Liu, Tulika Mitra
Pixel-space diffusion has recently emerged as a promising direction for high-fidelity image generation by modeling images directly in the original pixel domain. However, pixel-space diffusion is compu...
Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic errors such as missing...
TOLA is a diffusion‑based text image super‑resolution method that eliminates iterative image‑text diffusion by using a one‑step latent adaptation framework. It employs a confidence‑weighted text conditioning module to build a reliable semantic condition and a lightweight latent residual correction module to fix structured residual errors, thereby preserving text fidelity. Experiments show TOLA outperforms existing diffusion‑based TSR methods, achieving at least 2.72 dB higher PSNR on the CTR‑TSR‑Test benchmark.
By Yike Xu, Yue Shi, Yong Guo, Jiezhang Cao
arXiv:2608.29997v1 Announce Type: new
Abstract: We propose Discrete Diffusion Bridges (DDB), a novel framework designed to resolve the fundamental spatiotemporal misalignment of standard discrete dif...
By Xing Xie, Jiawei Liu, Shijun Zhou, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu
arXiv:2609.15120v1 Announce Type: new
Abstract: Benefiting from the powerful generative priors of diffusion models, diffusion-based real-world image super-resolution (Real-ISR) methods have demonstra...
By Shuhao Han, Wenjie Liao, Hayden Vance, Hang Dong, Rui Zhang, Chun-Le Guo, Chongyi Li
arXiv:2609.00798v1 Announce Type: new
Abstract: Pixel-space diffusion has recently emerged as a promising direction for high-fidelity image generation by modeling images directly in the original pixe...
By Weiyi You, Jinhua Zhang, Xingyu Zhou, Wei Long, Junyu Lou, Shuhang Gu
arXiv:2609.30988v1 Announce Type: new
Abstract: Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while...
By Xin Di, Mingyu Shi, Yuanfei Bao, Long Peng, Yue Zhao, Jiaming Guo, Renjing Pei, Xueyang Fu, Yang Cao, Zheng-Jun Zha