arXiv Machine Learning By Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak, Marco Fumero, Francesco Locatello, David Alvarez-Melis

Understanding Layer Patching in Model Size Interpolation

Read the original on arXiv Machine Learning →

arXiv:2607. 08170v1 Announce Type: new Abstract: Zero-shot model size interpolation aims to create new models of intermediate target sizes by combining existing models without additional training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
3d ago

GFD-OPD: Guidance-Folded On-Policy Distillation of Diffusion Models Across Scales

The paper introduces GFD-OPD, a method for on‑policy distillation of diffusion models that addresses challenges when compressing large teachers into smaller students. It identifies that standard distillation fails due to distribution gaps and classifier‑free guidance amplification, and proposes Fixed‑State KL to measure these gaps. GFD‑OPD reduces the student‑teacher discrepancy and achieves state‑of‑the‑art performance across multiple benchmarks.

By Zhenxing Zhang, Jiayan Teng, Wenxu Wu, Zhuoyi Yang, Jiazheng Xu, Wendi Zheng, Jie Tang, Dan Guo, Meng Wang