arXiv AI By Maksim Bobrin, Maksim Zhdanov, Dmitry Dylov

Fenchel Tilting: Weighted Correction for Efficient Finetuning of Generative Models

Read the original on arXiv AI →

Fenchel Tilt Flow Control (FTFC) is a new method for fine‑tuning pretrained generative models to arbitrary preference functions. It decouples utility optimization from model fitting by first learning reward and density‑ratio weights on pretrained samples, then freezing these weights to adjust a diffusion or flow model in a single importance‑weighted stage. The approach supports general f‑divergence penalties, achieves exact duality for concave utilities, and demonstrates up to 20× efficiency gains while outperforming baselines on image and molecule generation tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 24

WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps

The paper introduces Wasserstein‑Tilted Flow Maps (WTF), a simulation‑free reinforcement learning method that fine‑tunes pre‑trained flow‑based generative models by adding an optimal transport regularizer derived from the model’s drift. Unlike traditional KL‑reward tilting, WTF transports individual samples toward higher reward, framing the problem as a deterministic optimal control task on the flow map. Experiments on ImageNet‑256 and text‑to‑image demonstrate that WTF achieves higher reward and comparable or better diversity while reducing training compute by up to 280×.

By Abbas Mammadov, Jerry Y. Huang, Justin Lin, Partha Kaushik, Sheel Shah, Kartik Nair, Yee Whye Teh, Nicholas M. Boffi
arXiv Machine Learning
Jun 10

Exploring the Design Space of Reward Backpropagation for Flow Matching

arXiv:2606. 11075v1 Announce Type: new Abstract: Aligning text-to-image flow matching models with human preferences via direct reward backpropagation is sample-efficient but hampered by two well-known pathologies: activations cannot be stored across the full sampling trajectory at modern model scale, and chained Jacobian products across steps inflate the reward gradient as it travels back to early indices.

By Ruoyu Wang, Boye Niu, Xiangxin Zhou, Yushi Huang, Tongliang Liu, Chi Zhang