Self-OPD introduces a teacher‑free on‑policy distillation framework for flow matching models, using the student’s own exploration to generate step‑wise supervision. At each timestep the deterministic next‑state prediction is branched into multiple stochastic SDE candidates, rolled out, and compared against a deterministic baseline to compute normalized advantages. The velocity field is then optimized with a pull‑push objective that attracts high‑advantage branches and repels low‑advantage ones, while multi‑objective alignment is achieved by fusing normalized scores at the reward level.
By Shiyi Zhang, Mushui Liu, Yunze Tong, Wanggui He, Siyu Zou, Jinlong Liu, Yunlong Yu, Jian Song, Hao Jiang, Pipei Huang, Bo Zheng
The paper introduces GFD-OPD, a method for on‑policy distillation of diffusion models that addresses challenges when compressing large teachers into smaller students. It identifies that standard distillation fails due to distribution gaps and classifier‑free guidance amplification, and proposes Fixed‑State KL to measure these gaps. GFD‑OPD reduces the student‑teacher discrepancy and achieves state‑of‑the‑art performance across multiple benchmarks.
By Zhenxing Zhang, Jiayan Teng, Wenxu Wu, Zhuoyi Yang, Jiazheng Xu, Wendi Zheng, Jie Tang, Dan Guo, Meng Wang
The paper introduces Teacher Prediction Refinement Distillation (TPRD), a plug‑in module for Detection Transformers that refines teacher predictions before distillation. TPRD corrects degraded positive predictions and suppresses overconfident negatives, while preserving informative dark knowledge through Maximum Dark Knowledge Preservation. Experiments on MS COCO and PASCAL VOC show that these refinements improve the quality of supervision and the resulting student model’s performance.
By Yitong Xing, Yuhao Cheng, Yanping Li, Yichao Yan
arXiv:2608. 03632v1 Announce Type: new Abstract: On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals.
By Yinuo Jiang, Yongjie Ye, Zhou Tao, Xiang Zhuang, Qiang Zhang, Huajun Chen, Tiankai Li
DiffusionOPD introduces a multi-task training framework for diffusion models that leverages Online Policy Distillation (OPD). The method trains task-specific teachers separately and then distills their knowledge into a single student model using the student's own rollout trajectories, thereby separating exploration from integration. The authors extend OPD from discrete tokens to continuous-state Markov processes, deriving a closed-form per-step KL objective that unifies stochastic SDE and deterministic ODE refinement, and show that this analytic gradient yields lower variance and better generality than PPO-style gradients. Experiments demonstrate that DiffusionOPD outperforms both multi-reward RL and cascade RL baselines in training efficiency and final performance, achieving state-of-the-art results across all evaluated benchmarks.
By Quanhao Li, Junqiu Yu, Kaixun Jiang, Yujie Wei, Zhen Xing, Pandeng Li, Ruihang Chu, Shiwei Zhang, Yu Liu, Zuxuan Wu
Dataset distillation (DD) condenses large corpora into compact, information-rich subsets for efficient training and reuse. However, under noisy supervision, DD risks condensing corrupted associations together with useful signals, degrading robustness.