arXiv:2606. 09456v1 Announce Type: new Abstract: On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models.
By Yifan Niu, Han Xiao, Dongyi Liu, Zelong Wang, Dihong Gong, Yasheng Wang, Jia Li
arXiv:2606. 29869v1 Announce Type: cross Abstract: Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective often fail to balance primary distribution fitting with long-tail probability modeling, limiting both generation quality and generalization.
By Zilong Liu, Xuewen Zhang, Jinrui Xing, Juyi Qiao, Huiyong Wang, Junming Jiao
The paper introduces Tail‑Corrected Top‑k On‑Policy Distillation (TT‑OPD), a method that improves on existing Top‑k OPD by combining the selected top‑k tokens with a sampled token from the student’s rollout. This hybrid approach recovers the probability mass discarded by limiting to top‑k, yielding an unbiased estimator of the reverse KL divergence gradient while maintaining low computational cost. Experiments show TT‑OPD outperforms other OPD variants in accuracy.
By Linjian Meng, Siyuan Gan, YuHan Li, Xiran Wang, Ziyang Ding, Ditang Gou, Yiming Wu, Zhen Zhao
arXiv:2609.40235v1 Announce Type: cross
Abstract: Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (...
By Paul Le Van Kiem, Dario Shariatian, Umut Simsekli, Alain Durmus
Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective often fail to balance primary distribution fitting with long-tail probability modeling, limiting both generation quality and generalization. To address this, we analyze the complementary roles of forward and reverse KL divergence (FKL/RKL) in distribution alignment from theoretical and empirical perspectives.
arXiv:2610.07247v1 Announce Type: new
Abstract: Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for tr...
By Heng Liang, Xinwen Zhang, Hongchang Gao
arXiv:2603. 01875v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is an essential technique to compress large language models (LLMs) into smaller ones.
By Songming Zhang, Xue Zhang, Tong Zhang, Bojie Hu, Yufeng Chen, Jinan Xu
arXiv:2609.36734v1 Announce Type: new
Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implic...
By Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty
arXiv:2602. 22495v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs), but the high inference cost of such models motivates distillation into smaller students.
By Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang, Dhananjay Ram, Shuo Yang, Zhuowen Tu, Wei Xia, Stefano Soatto
arXiv:2402. 14035v4 Announce Type: replace-cross Abstract: Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality.
By Zichang Liu, Qingyun Liu, Yuening Li, Liang Liu, Anshumali Shrivastava, Shuchao Bi, Lichan Hong, Ed H. Chi, Zhe Zhao
arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.
By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
The paper introduces DLM-One, a score‑distillation framework that enables one‑step sequence generation with continuous diffusion language models (DLMs). By aligning a student model’s outputs with a pretrained teacher DLM’s score function in the forward‑diffused noisy space, DLM-One removes the need for iterative refinement. Experiments across various DLM architectures show up to ~2000× speedup in sampling steps and ~500× in wall‑clock time while retaining competitive performance, and the authors propose an adversarially‑regularized two‑stage training scheme to mitigate student degeneration.
By Tianqi Chen, Shujian Zhang, Mingyuan Zhou