arXiv AI By Jerry Y. Huang, Justin Lin, Sheel Shah, Kartik Nair, Nicholas M. Boffi

How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

Read the original on arXiv AI →

arXiv:2604. 27147v3 Announce Type: replace-cross Abstract: In generative modeling, we often wish to produce samples that maximize a user-specified reward such as aesthetic quality or alignment with human preferences, a problem known as \textit{guidance}.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 10

Exploring the Design Space of Reward Backpropagation for Flow Matching

arXiv:2606. 11075v1 Announce Type: new Abstract: Aligning text-to-image flow matching models with human preferences via direct reward backpropagation is sample-efficient but hampered by two well-known pathologies: activations cannot be stored across the full sampling trajectory at modern model scale, and chained Jacobian products across steps inflate the reward gradient as it travels back to early indices.

By Ruoyu Wang, Boye Niu, Xiangxin Zhou, Yushi Huang, Tongliang Liu, Chi Zhang
arXiv Machine Learning
Sep 24

WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps

The paper introduces Wasserstein‑Tilted Flow Maps (WTF), a simulation‑free reinforcement learning method that fine‑tunes pre‑trained flow‑based generative models by adding an optimal transport regularizer derived from the model’s drift. Unlike traditional KL‑reward tilting, WTF transports individual samples toward higher reward, framing the problem as a deterministic optimal control task on the flow map. Experiments on ImageNet‑256 and text‑to‑image demonstrate that WTF achieves higher reward and comparable or better diversity while reducing training compute by up to 280×.

By Abbas Mammadov, Jerry Y. Huang, Justin Lin, Partha Kaushik, Sheel Shah, Kartik Nair, Yee Whye Teh, Nicholas M. Boffi
arXiv AI
Sep 1

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies

CF‑VLA introduces a two‑stage coarse‑to‑fine approach for vision‑language‑action policies, replacing multi‑step sampling with a coarse initialization that constructs an action‑aware starting point and a single‑step refinement that corrects residual errors. The coarse stage learns a conditional posterior over endpoint velocity to transform Gaussian noise into a structured initialization, while the fine stage performs a fixed‑time refinement. Experiments on CALVIN and LIBERO demonstrate that CF‑VLA achieves a strong efficiency‑performance trade‑off, reducing action sampling latency by 75.4 % and achieving an 83.0 % real‑robot success rate, outperforming existing NFE=2 methods and matching or surpassing NFE=10 baselines.

By Fan Du, Feng Yan, Jianxiong Wu, Xinrun Xu, Weiye Zhang, Weinong Wang, Yu Guo, Bin Qian, Zhihai He, Fei Wang, Heng Yang