arXiv Computer Vision By Anne Gagneux, S\'egol\`ene Martin, R\'emi Gribonval, Mathurin Massias

Training Flow Matching: The Role of Weighting and Parameterization

Read the original on arXiv Computer Vision →

The paper investigates training objectives for denoising-based generative models, focusing on loss weighting and output parameterization such as noise-, clean image-, and velocity-based formulations. It conducts a systematic numerical study across synthetic datasets with controlled geometry and real image data, evaluating denoising accuracy via PSNR and generative quality via FID. The goal is to disentangle how training choices interact with data manifold dimensionality, model architecture, and dataset size, offering practical design insights rather than proposing a new method.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Jul 13

Self-Consistent Flow: Unifying Velocity and Endpoint Prediction for Rectified Flow Models

In rectified-flow-based generative models, the neural network can be trained to predict two different targets, such as the instantaneous velocity or the data endpoint, to perform denoising. Although prior work shows that these parameterizations lead to different empirical behaviors, the mechanisms underlying their respective advantages remain to be underexplored, and how to combine them effectively is still unclear.

arXiv Machine Learning
Sep 24

On the Diffusibility of High-Dimensional Latents

The paper investigates how fine‑tuning pretrained visual encoders for faithful image reconstruction affects diffusion models that operate in the resulting latent space. It finds that such fine‑tuning reduces the effective dimensionality of the latent representation, causing standard velocity‑prediction flow‑matching to fit noise outside the low‑dimensional signal manifold and making optimization inefficient. Consequently, the authors propose using a clean‑data ($oldsymbol{x}_{0}$) parameterization, which focuses learning on the signal manifold and consistently improves text‑to‑image generation across multiple strong‑reconstruction encoders.

By Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun Xiong, Xiyao Wang, Jui-Hsien Wang, Richard Zhang, Zhe Lin, Andrew Owens, Yijun Li