arXiv:2606. 29724v1 Announce Type: new Abstract: Normalizing flows are powerful generative models that learn an invertible mapping between complex data distributions and simple latent distributions, typically a standard normal density.
By Liam A. Kruse, Houjun Liu, Alexandros E. Tzikas, Mansur M. Arief, Mykel J. Kochenderfer
Normalizing flows are powerful generative models that learn an invertible mapping between complex data distributions and simple latent distributions, typically a standard normal density. However, this choice of latent density can impose unnecessary complexity on the learned flow transformation due to the topological mismatch between the latent and data densities, leading to slower training and suboptimal performance.
arXiv:2605.06272v2 Announce Type: replace
Abstract: While generative modeling has achieved remarkable success on tasks like natural language-conditioned image generation, enabling model adaptation fr...
By Tyler Ingebrand, Ruihan Zhao, Kushagra Gupta, David Fridovich-Keil, Sandeep P. Chinchali, Ufuk Topcu
arXiv:2605. 08398v2 Announce Type: replace Abstract: In this work, we show that Latent Flow-Matching (LFM) models are robust to different types of perturbations, including data reduction and model capacity shrinkage.
By Rania Briq, Michael Kamp, Ohad Fried, Sarel Cohen, Stefan Kesselheim
Geometric foundation models, such as the Visual Geometry Grounded Transformer (VGGT), provide strong 3D priors from unposed images. However, such models operate purely in a feed-forward, deterministic regime, \ie~they cannot generate plausible geometry beyond what the input views directly support.
arXiv:2609.35763v3 Announce Type: replace
Abstract: Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representa...
By Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu
arXiv:2602. 05951v2 Announce Type: replace-cross Abstract: Flow matching has recently emerged as a promising alternative to diffusion-based generative models, particularly for text-to-image generation.
By Junwan Kim, Jiho Park, Seonghu Jeon, Seungryong Kim
arXiv:2607. 23946v1 Announce Type: new Abstract: We introduce Joint Flow Matching (JFM), a training framework for continuous normalising flows over multiple variables.
By Hayden McAlister, Lech Szymanski
The paper introduces continuous adversarial flow models, a continuous-time flow framework trained with an adversarial objective that replaces the fixed mean-squared-error criterion of flow matching. By incorporating a learned discriminator, the method guides training toward a different generalized distribution, yielding samples more closely aligned with the target data distribution. Applied as a post‑training step, it markedly improves ImageNet 256px generation metrics—reducing the guidance‑free FID of latent‑space SiT from 8.26 to 3.63 and of pixel‑space JiT from 7.17 to 3.57—and also enhances guided generation and text‑to‑image benchmarks.
By Shanchuan Lin, Ceyuan Yang, Zhijie Lin, Hao Chen, Haoqi Fan
arXiv:2608. 08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task.
By Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
The paper introduces GeoNeXt, a framework that repurposes pretrained video generative models for geometry estimation by framing it as a next‑frame prediction task. Unlike prior methods that either train separate depth/normal models or fine‑tune image diffusion backbones, GeoNeXt jointly models images and geometric targets, leveraging the structured knowledge of video models for more data‑efficient learning. Experiments show zero‑shot monocular depth and surface normal estimation that outperforms existing generative approaches and rivals discriminative state‑of‑the‑art methods while using far less training data.
By Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
The paper investigates how fine‑tuning pretrained visual encoders for faithful image reconstruction affects diffusion models that operate in the resulting latent space. It finds that such fine‑tuning reduces the effective dimensionality of the latent representation, causing standard velocity‑prediction flow‑matching to fit noise outside the low‑dimensional signal manifold and making optimization inefficient. Consequently, the authors propose using a clean‑data ($oldsymbol{x}_{0}$) parameterization, which focuses learning on the signal manifold and consistently improves text‑to‑image generation across multiple strong‑reconstruction encoders.
By Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun Xiong, Xiyao Wang, Jui-Hsien Wang, Richard Zhang, Zhe Lin, Andrew Owens, Yijun Li