arXiv AI

MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators

Hugging Face Trending Papers
Aug 20

Continuous Adversarial MeanFlow Transfer

Training fast generators on new domains with limited data remains challenging for two reasons. First, adapting a pretrained diffusion or flow model to a new domain leaves its costly multi-step sampling unaddressed, and existing acceleration methods are tied to the source parameterization--$ε$, $x$, $v$, or $u$--leaving heterogeneous pretrained models with no common acceleration target.

arXiv Machine Learning
Sep 24

WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps

The paper introduces Wasserstein‑Tilted Flow Maps (WTF), a simulation‑free reinforcement learning method that fine‑tunes pre‑trained flow‑based generative models by adding an optimal transport regularizer derived from the model’s drift. Unlike traditional KL‑reward tilting, WTF transports individual samples toward higher reward, framing the problem as a deterministic optimal control task on the flow map. Experiments on ImageNet‑256 and text‑to‑image demonstrate that WTF achieves higher reward and comparable or better diversity while reducing training compute by up to 280×.

By Abbas Mammadov, Jerry Y. Huang, Justin Lin, Partha Kaushik, Sheel Shah, Kartik Nair, Yee Whye Teh, Nicholas M. Boffi
arXiv AI
Sep 4

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.

By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv Computer Vision
Aug 24

Difficulty-Calibrated Interpolation Paths for Conditional Flow Matching

The paper introduces Difficulty-Calibrated Flow Matching, a method that adapts the noise-to-data interpolation schedule in Conditional Flow Matching based on a pilot run’s loss profile. By setting the schedule to the quantile function of this difficulty profile, the training trajectory spends more time where the velocity is hardest to learn. Experiments on CIFAR-10, MNIST, and Fashion‑MNIST show that this calibrated path achieves the best FID on CIFAR‑10 and outperforms all fixed schedules in large‑batch, few‑update settings, where compute is most limited.

By Airin Akter Tania, Md Raihan Khan