arXiv:2607. 22725v1 Announce Type: cross Abstract: Data augmentation is routinely used to improve generalization in image classification, but the assumptions underlying standard policies are poorly matched to coherent imaging.
By Mohamed Abdallah Salem, Nourhan Zein Diab
arXiv:2607. 05319v1 Announce Type: cross Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures.
By Rajat Rasal, Avinash Kori, Tian Xia, Ben Glocker
arXiv:2606. 02661v1 Announce Type: cross Abstract: Accurate precipitation nowcasting is vital for disaster mitigation, but deep learning methods face a key trade-off: regression models produce over-smoothed, spectrally decaying predictions that blur convective details and violate turbulence power laws; diffusion models generate realistic yet unanchored hallucinations lacking physical grounding.
By Yunlong Zhou, Chen Zhao, Danyang Peng, Fanfan Ji, Xiao-Tong Yuan
arXiv:2607. 10140v1 Announce Type: cross Abstract: Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation.
By Yuang Meng, Chenyang Wu, Xianshun Liu, Chun-Le Guo, Zichen Liang, Lina Lei, Jie Liang, Hui Zeng, Chongyi Li, Lei Zhang
arXiv:2606. 11691v1 Announce Type: new Abstract: Latent diffusion and flow matching have emerged as leading approaches for synthetic turbulence generation, yet they systematically under-represent dissipation-range amplitudes.
By Khalid Rafiq, Aditya G. Nair
The paper introduces a perceptually regularized diffusion framework for image super‑resolution, adding perceptual‑loss based regularization to the standard diffusion training objective. This approach incorporates prior knowledge to improve training convergence and encourages the recovery of meaningful image features. Experiments on benchmark datasets show enhanced perceptual quality while maintaining competitive distortion metrics.
By Chuxiangbo Wang, Pavithra Venkatachalapathy, Ying Liang, Min Wang, Jing Qin, Yifei Lou, Weihong Guo
arXiv:2609.38438v1 Announce Type: cross
Abstract: Machine learning surrogate models offer a promising path toward accelerating plasma turbulence simulations. We present PreVAE-Turb, a surrogate model...
By Minglei Yang, Marshall Nicholson, Diego Del-Castillo-Negrete, David Hatch, Guannan Zhang
Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full se...
arXiv:2607. 12464v1 Announce Type: cross Abstract: When labeled data are scarce, off-the-shelf diffusion models can augment training sets for few-shot medical image classification, but not all generated samples are equally useful for the downstream task.
By Jeeyung Kim, Erfan Esmaeili, Qiang Qiu
arXiv:2507. 00719v3 Announce Type: replace-cross Abstract: Typically, numerical simulations of Earth systems are coarse, and Earth observations are sparse and gappy.
By Anantha Narayanan Suresh Babu, Akhil Sadam, Pierre F. J. Lermusiaux
FastVR is a streaming video restoration framework that uses a one‑step diffusion model to achieve strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. It addresses efficiency bottlenecks by combining a lightweight VAE with chunk‑wise causal attention, and improves inference speed and restoration quality through velocity consistency regularization and continuous trajectory learning during training. Experiments demonstrate that FastVR outperforms diffusion baselines in efficiency and achieves state‑of‑the‑art performance on both synthetic and real‑world benchmarks.
By Xiaoxu Chen, Qin Yang, Haoran Bai, Sibin Deng, Ying Chen
Pixel‑Space Diffusion via Observation Operators introduces a new framework for pixel‑space diffusion models that addresses a scale‑time mismatch in existing methods. By replacing fixed full‑image supervision with a time‑indexed observation trajectory that progresses from coarse structures to the full image, the model aligns supervision with the natural recovery order of image details. The approach employs Gaussian‑Lanczos operators and a GL‑CoDA decoder to refine features progressively, resulting in faster convergence and higher generation quality, achieving an FID of 1.52 on ImageNet‑256.
By Shaojie Guo, Lichen Ma, Haoyang Tong, Yu He, Zipeng Guo, Xiaoan Liu, Feng Yan, Yu Guo, Fei Wang, Junshi Huang, Yan Wang