The paper introduces a Focal Log-Frequency Loss (f-loss) to counteract the spectral imbalance in pixel-space flow matching, where low frequencies dominate training. By balancing learning signals across frequencies and combining early frequency-domain supervision with later pixel-space refinement, the method accelerates convergence by up to 40% and improves FID and perceptual fidelity across multiple model scales. It requires no architectural changes and can replace existing flow matching losses as a drop‑in solution.
By Lucas Degeorge, Paul Couairon, Arijit Ghosh, Alexei A. Efros, David Picard, Vicky Kalogeiton
The paper introduces continuous adversarial flow models, a continuous-time flow framework trained with an adversarial objective that replaces the fixed mean-squared-error criterion of flow matching. By incorporating a learned discriminator, the method guides training toward a different generalized distribution, yielding samples more closely aligned with the target data distribution. Applied as a post‑training step, it markedly improves ImageNet 256px generation metrics—reducing the guidance‑free FID of latent‑space SiT from 8.26 to 3.63 and of pixel‑space JiT from 7.17 to 3.57—and also enhances guided generation and text‑to‑image benchmarks.
By Shanchuan Lin, Ceyuan Yang, Zhijie Lin, Hao Chen, Haoqi Fan
arXiv:2607. 28760v1 Announce Type: cross Abstract: As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality.
By Krunoslav Lehman Pavasovic, Th\'eophane Vallaeys, St\'ephane Mallat, Giulio Biroli, Luke Zettlemoyer, Brian Karrer, Jakob Verbeek
arXiv:2603.20186v2 Announce Type: replace
Abstract: In this work, we propose Image-to-Image Rectified Flow Reformulation (I2I-RFR), a practical plug-in reformulation that recasts standard I2I regress...
By Satoshi Iizuka, Shun Okamoto, Kazuhiro Fukui
arXiv:2609.35763v3 Announce Type: replace
Abstract: Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representa...
By Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu
arXiv:2607. 12171v1 Announce Type: cross Abstract: In rectified-flow-based generative models, the neural network can be trained to predict two different targets, such as the instantaneous velocity or the data endpoint, to perform denoising.
By Xu Han, Jiajing Hu, Li-Ping Liu