arXiv:2607. 17572v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences.
By Ruiyi Ding, Jie Li, He Kang, Ziyan Liu, Chengru Song, Yuan chen
arXiv:2607. 17572v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences.
By Ruiyi Ding, Jie Li, He Kang, Ziyan Liu, Chengru Song, Yuan chen
The paper introduces the Latent Generative Solver (LGS), a neural PDE solver that combines a Physics VAE, a Pyramidal Flow-Forcing Transformer, and input noising to achieve generalization across twelve PDE families and stable long-term rollouts. LGS matches or surpasses deterministic baselines on one-step predictions, outperforms them on 5- and 10-step rollouts, and significantly reduces long-horizon error while cutting compute costs. It also adapts efficiently to unseen higher-resolution systems, demonstrating strong empirical performance on 2D regular-grid PDE simulations.
By Zituo Chen, Sili Deng
The paper introduces Kinematic MeanFlow (K-MF), a one‑step action generation policy for Robotic Foundation Models that addresses instability in the MeanFlow framework. By decoupling the time derivative into two sub‑interval terms, K-MF captures early and late denoising dynamics separately, reducing error amplification. Experiments show K-MF achieves faster inference—reducing action‑head latency by 67.5%–74.4% and overall end‑to‑end latency by 30.3%–54.9%—while outperforming multi‑step flow matching on various tasks.
By Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao
Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences. While successful in large language models~\cite{shao2024deepseekmathpushinglimitsmathematical}, its extension to diffusion and flow matching models introduces a severe computational bottleneck: gradients must be back-propagated through the high-capacity DiT backbone at \emph{every} timestep of the sampling trajectory, making high-resolution text-to-image (T2I) training prohibitively expensive.
arXiv:2607. 23667v1 Announce Type: cross Abstract: A flow surrogate validated on a simple regime is often taken as evidence that the approach will carry to a richer one.
By Georg Winkler, Martin Stoll
arXiv:2607. 21644v1 Announce Type: new Abstract: We present a goal-agnostic control framework for partial differential equations (PDEs) built around a joint-embedding predictive architecture (JEPA).
By Jonathan Gallagher, Roberto Guglielmi
The paper introduces TDAction, a method for one‑step generative modeling that selects transport targets during training based on a cost reflecting shared‑parameter effort and terminal mismatch. By formulating this as a soft‑terminal control problem, the authors derive a closed‑form Batch Tangent Action‑to‑Go value that captures cross‑sample interactions and can be efficiently implemented with randomized tangent probes. Experiments on ImageNet 256×256 demonstrate that TDAction achieves an FID below 1.1 without distillation.
By Zhangyong Liang, Ying Huang, Haibin Ling
In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a pr...
arXiv:2602.12624v2 Announce Type: replace
Abstract: Diffusion-based generative models have achieved remarkable performance across various domains, yet their practical deployment is often limited by h...
By Sangwoo Jo, Sungjoon Choi
arXiv:2608. 19540v1 Announce Type: new Abstract: Training fast generators on new domains with limited data remains challenging for two reasons.
By Yara Bahram, Zahra Dehghani, M\'elodie Desbos, Eric Granger, Pablo Piantanida, Mohammadhadi Shateri
arXiv:2607. 11442v1 Announce Type: new Abstract: Flow matching trains a neural network to regress the conditional velocity along a linear interpolant between noise and data, and the number of network evaluations~(NFE) sets the cost of sampling.
By Vitalii Bondar