arXiv Machine Learning By Tiancong Cheng, Ying Zhang, Zhiwen Yu, Yifang Yin, Bin Guo

Progressive$^2$: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression

Read the original on arXiv Machine Learning →

arXiv:2608. 00129v1 Announce Type: new Abstract: Knowledge distillation (KD) is a widely utilized technique for transferring knowledge from a large model (the teacher) to a smaller model (the student).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.