arXiv:2609.09905v1 Announce Type: cross
Abstract: Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation betw...
By Yansen Han, Shengyi Liao, Peng Sun, Deyuan Liu, Yuanxing Zhang, Pengfei Wan, Tao Lin
arXiv:2601. 22495v2 Announce Type: replace Abstract: Fine-tuning flow matching models is a central challenge in settings with limited data, evolving distributions, or computational constraints.
By Gudrun Thorkelsdottir, Arindam Banerjee
arXiv:2606. 11075v1 Announce Type: new Abstract: Aligning text-to-image flow matching models with human preferences via direct reward backpropagation is sample-efficient but hampered by two well-known pathologies: activations cannot be stored across the full sampling trajectory at modern model scale, and chained Jacobian products across steps inflate the reward gradient as it travels back to early indices.
By Ruoyu Wang, Boye Niu, Xiangxin Zhou, Yushi Huang, Tongliang Liu, Chi Zhang
arXiv:2604. 17415v3 Announce Type: replace-cross Abstract: Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model.
By Jeongjae Lee, Jinho Chang, Jeongsol Kim, Jong Chul Ye
arXiv:2607. 10853v1 Announce Type: cross Abstract: Diffusion models faithfully reproduce their training distribution, but also inherit its imbalances and leave rare or under-represented modes hard to reach.
By Peizhuo Li, Emre Aksan, Alexandru-Eugen Ichim, Thabo Beeler, Olga Sorkine-Hornung
arXiv:2609.05727v1 Announce Type: cross
Abstract: We develop Newton Matching, a unified framework for fine-tuning and sampling in generative modeling. The target is $\pi\propto\mu e^{\tau r}$, where...
By Zeyang Li, Yunan Wang, Paolo Giaretta, Navid Azizan
arXiv:2607. 08781v1 Announce Type: cross Abstract: The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice.
By Kehan Guo, Yili Shen, Yujun Zhou, Yue Huang, Chujie Gao, Shiyi Du, Xiangliang Zhang
arXiv:2606. 02521v1 Announce Type: new Abstract: One-step text-to-image generators are attractive for deployment because they generate an image with a single forward pass, but preference finetuning them remains difficult: standard alignment methods often rely on policy likelihoods, denoising trajectories, differentiable reward gradients, or test-time optimization.
By Zhou Jiang, Yandong Wen, Zhen Liu
arXiv:2604. 27147v3 Announce Type: replace-cross Abstract: In generative modeling, we often wish to produce samples that maximize a user-specified reward such as aesthetic quality or alignment with human preferences, a problem known as \textit{guidance}.
By Jerry Y. Huang, Justin Lin, Sheel Shah, Kartik Nair, Nicholas M. Boffi
The paper examines the common practice in diffusion and flow matching of predicting the clean signal, converting it to a velocity, and training with a velocity-space loss. It shows that this conversion can cause unstable optimization due to singular endpoint amplification, but that prediction–loss alignment removes this source of non‑integrability and ensures a finite second moment for all timesteps, even with uniform sampling. Experiments confirm that aligned objectives remain trainable across different samplers, reconciling theoretical concerns with empirical success.
By Jiadong Hong, Lei Liu, Xinyu Bian, Wenjie Wang, Zhaoyang Zhang
arXiv:2607. 00947v1 Announce Type: new Abstract: Generative models learn data distributions that reside on a low-dimensional manifold within a higher-dimensional ambient space.
By Ludwig Winkler, Andrew Leaver-Fay, Joseph Kleinhenz, Pan Kessel
arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.
By Hyung Gyu Rho