The paper introduces Flow Divergence Sampler (FDS), a training‑free method that refines intermediate states in flow‑matching models by using the divergence of the marginal velocity field to detect and correct misguidance toward low‑density regions. FDS operates during inference, requires no additional training, and can be applied as a plug‑and‑play module with standard solvers and existing flow backbones. Experiments show that FDS consistently improves fidelity in tasks such as text‑to‑image synthesis and inverse problems.
By Yeonwoo Cha, Jaehoon Yoo, Semin Kim, Yunseo Park, Jinhyeon Kwon, Seunghoon Hong
The paper studies classifier‑free guidance (CFG) in Flow Matching, showing that strong guidance can distort the generated distribution by shifting the mean and concentrating trajectories. By interpreting Flow Matching as a time‑varying gradient flow, the authors explain how CFG reshapes the underlying potential and propose a training‑free method, Posterior‑Mean‑Capped CFG (PMC‑CFG), that adaptively limits guidance to the strongest feasible level. Experiments on synthetic and large‑scale image‑generation tasks demonstrate that PMC‑CFG reduces distortion and concentration while improving the alignment–diversity trade‑off, especially when nominal guidance is large.
By Jishen Peng, Zheng Ma
arXiv:2608.29107v1 Announce Type: new
Abstract: While modern generative models excel at modeling complex data, precise inference-time control in conditional generation remains a critical challenge. C...
By Avishag Nevo, Tamir Hazan
arXiv:2502. 08006v3 Announce Type: replace-cross Abstract: Training-free guided generation is a widely used and powerful technique that allows the end user to exert further control over the generative process of flow/diffusion models.
By Zander W. Blasingame, Chen Liu
arXiv:2607. 12171v1 Announce Type: cross Abstract: In rectified-flow-based generative models, the neural network can be trained to predict two different targets, such as the instantaneous velocity or the data endpoint, to perform denoising.
By Xu Han, Jiajing Hu, Li-Ping Liu
arXiv:2606. 27377v1 Announce Type: cross Abstract: Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing.
By Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, Lixue Gong, Yongyuan Liang, Meng Chu, Leigang Qu, Lingdong Kong, Wei Liu, Tat-Seng Chua
Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly.
arXiv:2607. 09133v1 Announce Type: cross Abstract: While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solvers incurs severe inference latency.
By Yiting Wang, Jingyi Zhang, Wenhu Zhang, Ke Chao, Yves Liang, Kun Cheng, Kang Zhao
arXiv:2601. 23231v2 Announce Type: replace-cross Abstract: Flow-based generative models provide strong unconditional priors for inverse problems, but guiding their dynamics for conditional generation remains challenging.
By George Webber, Alexander Denker, Riccardo Barbano, Andrew J Reader
arXiv:2604. 27147v3 Announce Type: replace-cross Abstract: In generative modeling, we often wish to produce samples that maximize a user-specified reward such as aesthetic quality or alignment with human preferences, a problem known as \textit{guidance}.
By Jerry Y. Huang, Justin Lin, Sheel Shah, Kartik Nair, Nicholas M. Boffi
arXiv:2607. 26398v1 Announce Type: new Abstract: Diffusion and flow-based models benefit from simple regression losses, but inference incurs significant overhead because sampling requires integration.
By Mark Goldstein, Anshuk Uppal, Raghav Singhal, Aahlad Puli, Rajesh Ranganath
In rectified-flow-based generative models, the neural network can be trained to predict two different targets, such as the instantaneous velocity or the data endpoint, to perform denoising. Although prior work shows that these parameterizations lead to different empirical behaviors, the mechanisms underlying their respective advantages remain to be underexplored, and how to combine them effectively is still unclear.