arXiv Machine Learning By Mingfeng Lin, Chengfei Cai, Lin Xu, Yuxiang Wei, Liang Han

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models

Read the original on arXiv Machine Learning →

arXiv:2608. 09233v1 Announce Type: new Abstract: Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Aug 28

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

Self-OPD introduces a teacher‑free on‑policy distillation framework for flow matching models, using the student’s own exploration to generate step‑wise supervision. At each timestep the deterministic next‑state prediction is branched into multiple stochastic SDE candidates, rolled out, and compared against a deterministic baseline to compute normalized advantages. The velocity field is then optimized with a pull‑push objective that attracts high‑advantage branches and repels low‑advantage ones, while multi‑objective alignment is achieved by fusing normalized scores at the reward level.

By Shiyi Zhang, Mushui Liu, Yunze Tong, Wanggui He, Siyu Zou, Jinlong Liu, Yunlong Yu, Jian Song, Hao Jiang, Pipei Huang, Bo Zheng
arXiv Machine Learning
Sep 14

DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

DiffusionOPD introduces a multi-task training framework for diffusion models that leverages Online Policy Distillation (OPD). The method trains task-specific teachers separately and then distills their knowledge into a single student model using the student's own rollout trajectories, thereby separating exploration from integration. The authors extend OPD from discrete tokens to continuous-state Markov processes, deriving a closed-form per-step KL objective that unifies stochastic SDE and deterministic ODE refinement, and show that this analytic gradient yields lower variance and better generality than PPO-style gradients. Experiments demonstrate that DiffusionOPD outperforms both multi-reward RL and cascade RL baselines in training efficiency and final performance, achieving state-of-the-art results across all evaluated benchmarks.

By Quanhao Li, Junqiu Yu, Kaixun Jiang, Yujie Wei, Zhen Xing, Pandeng Li, Ruihang Chu, Shiwei Zhang, Yu Liu, Zuxuan Wu
arXiv Machine Learning
Jul 28

FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

arXiv:2607. 24522v1 Announce Type: new Abstract: While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored.

By Kaiyang Ye, Yuan Ge, Junxiang Zhang, Bei Li, Ziming Zhu, Haishu Zhao, Xiaoqian Liu, Chenglong Wang, Jingbo Zhu, Zhengtao Yu, Tong Xiao
Hugging Face Trending Papers
Aug 10

RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We instead take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate states provide a natural source of distillation supervision rather than a disposable byproduct of sampling.

arXiv AI
6d ago

DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

DyMD introduces a Distribution Matching Distillation framework that adapts teacher supervision and critic fitting to preserve interaction dynamics in few-step video generation. By employing temporal affinity–conditioned re‑noise sampling and dynamics‑guided fake‑score tracking, DyMD balances motion recovery with visual quality. The method distills a 14B teacher into a 1.3B student that achieves significant gains on embodied‑video benchmarks and downstream action planning tasks.

By Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu