arXiv:2606. 16742v1 Announce Type: cross Abstract: With the rapid advancement of video generation models, distinguishing between AI-generated and authentic videos has emerged as a challenging endeavor.
By Renxi Cheng, Jie Gui, Hongsong Wang
arXiv:2404. 06294v2 Announce Type: replace-cross Abstract: Super-Resolution (SR) is a time-hallowed image processing problem that aims to improve the quality of a Low-Resolution (LR) sample up to the standard of its High-Resolution (HR) counterpart.
By Arkaprabha Basu, Kushal Bose, Sankha Subhra Mullick, Anish Chakrabarty, Swagatam Das
arXiv:2608. 14706v1 Announce Type: cross Abstract: Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampling schedules, limiting inference procedures from adapting to the data.
By Hansen Jin Lillemark, Alex Rojas, Zachary Novack, Runqian Wang, Yilun Du, Yian Ma, Taylor Berg-Kirkpatrick, Rose Yu
arXiv:2607. 19332v1 Announce Type: new Abstract: Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching.
By Chirag Vashist, Ke Li
The paper introduces Contrastive Noise Alignment (CNA), a training-time method for generative flow models that dynamically aligns Gaussian noise with data samples using a cross-modal InfoNCE objective. By modeling noise as an interacting particle system and regularizing with angular entropy and radial norm penalties, CNA reduces arbitrary data-noise couplings and flow curvature. Empirical results show that CNA improves generation quality, lowering FID by over 50% for few-step pixel-space generation compared to standard rectified flow and outperforming optimal transport baselines by at least 24%.
By Lennart Wittke, Vinicius Azevedo
In rectified-flow-based generative models, the neural network can be trained to predict two different targets, such as the instantaneous velocity or the data endpoint, to perform denoising. Although prior work shows that these parameterizations lead to different empirical behaviors, the mechanisms underlying their respective advantages remain to be underexplored, and how to combine them effectively is still unclear.
arXiv:2607. 12171v1 Announce Type: cross Abstract: In rectified-flow-based generative models, the neural network can be trained to predict two different targets, such as the instantaneous velocity or the data endpoint, to perform denoising.
By Xu Han, Jiajing Hu, Li-Ping Liu
Pixel diffusion models generate RGB images directly but tend to miss fine‑scale natural‑image statistics. The authors introduce an adversarial post‑training step that adds an adversarial loss to the model’s output at non‑high‑noise timesteps, without changing the architecture or sampling procedure. This approach improves distribution fidelity, coverage, prompt alignment, and perceptual quality across two pixel backbones, and restores missing high‑frequency spectral power while avoiding memorization or mode dropping.
By Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen
arXiv:2608. 10544v1 Announce Type: cross Abstract: Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations.
By Sangwoo Jo, Donggeun Ko, Jayeon Kang, Youngsang Kwak, Jaehwa Kwak, Sungjoon Choi
Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching. Along the way, the underlying techniques have become more complicated and various beliefs about what drives strong empirical performance have taken hold.
arXiv:2605. 23264v2 Announce Type: replace-cross Abstract: Generative priors in Image Super-Resolution (SR) often compromise faithful restoration, we attribute this limitation to a fundamental spectral misalignment between isotropic objectives and the intrinsic natural image manifold.
By Hongbo Wang, Huaibo Huang, Pin Wang, Jinhua Hao, Chao Zhou, Ran He
The paper investigates training objectives for denoising-based generative models, focusing on loss weighting and output parameterization such as noise-, clean image-, and velocity-based formulations. It conducts a systematic numerical study across synthetic datasets with controlled geometry and real image data, evaluating denoising accuracy via PSNR and generative quality via FID. The goal is to disentangle how training choices interact with data manifold dimensionality, model architecture, and dataset size, offering practical design insights rather than proposing a new method.
By Anne Gagneux, S\'egol\`ene Martin, R\'emi Gribonval, Mathurin Massias