arXiv Machine Learning

Why Are DMD Students Lazy? Understanding the Copying Behavior in Few-Step Distillation

arXiv:2606. 02237v1 Announce Type: new Abstract: Distribution Matching Distillation (DMD) compresses pretrained diffusion models into efficient few-step generators by aligning their noised distributions across all scales.

arXiv AI
3d ago

GFD-OPD: Guidance-Folded On-Policy Distillation of Diffusion Models Across Scales

The paper introduces GFD-OPD, a method for on‑policy distillation of diffusion models that addresses challenges when compressing large teachers into smaller students. It identifies that standard distillation fails due to distribution gaps and classifier‑free guidance amplification, and proposes Fixed‑State KL to measure these gaps. GFD‑OPD reduces the student‑teacher discrepancy and achieves state‑of‑the‑art performance across multiple benchmarks.

By Zhenxing Zhang, Jiayan Teng, Wenxu Wu, Zhuoyi Yang, Jiazheng Xu, Wendi Zheng, Jie Tang, Dan Guo, Meng Wang
arXiv Machine Learning
Jun 5

OPRD: On-Policy Representation Distillation

arXiv:2606. 06021v1 Announce Type: new Abstract: On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities.

By Shenzhi Yang, Guangcheng Zhu, Bowen Song, Haobo Wang, Mingxuan Xia, Xing Zheng, Yingfan Ma, Zhongqi Chen, Weiqiang Wang, Gang Chen
Hugging Face Trending Papers
Aug 4

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language: identical VAE latents, matching architectures, and a common timestep grid. We ask what happens when none of this holds, as when the strongest teacher available and the student one wishes to deploy come from different model families, and find that the standard recipes have no answer: teacher latents cannot serve as targets in a foreign coordinate system, per-pixel losses against a teacher that stochastically re-draws local detail degenerate into blur or divergence, and timestep indices lose their meaning across mismatched schedules.

arXiv Computer Vision
Aug 26

IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves

The paper introduces IDeaL, a data‑free multi‑teacher distillation technique that generates teacher‑specific, improved samples using decorrelation losses at patch and image levels. By tailoring noise to each teacher, IDeaL produces strong student models that capture complementary teacher information and achieve results close to those distilled from real images. Experiments demonstrate that with only 1,000 images, students trained on IDeaL samples match or exceed the performance of students distilled from a 1,000‑image subset of ImageNet.

By Feyza Yavuz, Mert B\"ulent Sar{\i}y{\i}ld{\i}z, Diane Larlus
arXiv Machine Learning
Aug 5

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

arXiv:2608. 03316v1 Announce Type: new Abstract: On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language: identical VAE latents, matching architectures, and a common timestep grid.

By Siming Fu, Zheming Fu, Ruizhe He, Hualiang Wang, Jie Huang, Xiaoxiao Ma, Mingchen Zhong, Weihu Huang, Xiaoxuan He, Haojun Xu