arXiv:2609.37500v1 Announce Type: new
Abstract: On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its relianc...
By Yuxiao Yang, Shangzhe Li, Tianrun Yu, Kaixiang Zhao, Taylor W. Killian, Weitong Zhang
Activation-Conditioned Self-Distillation (ACSD) is a new on‑policy self‑distillation method that uses a frozen copy of the base model to extract a steering vector by contrasting activations from self‑generated trajectories that reach verified correct answers with all other trajectories. The student learns from next‑token distributions on its own prefixes, without needing reference text or teacher parameter updates, and is used alone at inference. Across five models, ACSD achieves the highest mean accuracy on four mathematical benchmarks, with notable gains on DeepSeek‑R1‑0528‑Qwen3‑8B and LiveCodeBench v6 compared to the OPSD baseline.
By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv:2604. 05634v2 Announce Type: replace Abstract: Machine unlearning (MU) has become a critical technique for GenAI models' safe and compliant operation.
By Zhiyong Ma, Zhitao Deng, Huan Tang, Jialin Chen, Zhijun Zheng, Zhengping Li, Qingyuan Chuai
arXiv:2609.36695v1 Announce Type: new
Abstract: Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external tea...
By Rui Wang, Ruijie Wang, Bo Chen, Jiangxuan Long, Yingyu Liang
arXiv:2606. 09456v1 Announce Type: new Abstract: On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models.
By Yifan Niu, Han Xiao, Dongyi Liu, Zelong Wang, Dihong Gong, Yasheng Wang, Jia Li
VISTA is an online self‑distillation framework that enforces consistency along a deep learning model’s optimization trajectory. It uses a validation‑informed Marginal Coverage score to identify earlier model states—called expert anchors—that retain specialized competence over distinct data regions. By integrating a coverage‑weighted ensemble of these anchors during training, VISTA regularizes the loss landscape, preserves learned knowledge, and improves robustness and generalization while cutting storage overhead by 90%.
By Eli Corn, Daphna Weinshall
arXiv:2605.10194v2 Announce Type: replace
Abstract: On-policy self-distillation (OPSD) uses a model as its own teacher under privileged context, providing token-level supervision on the model's own r...
By Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen, Yulan Hu, Lan-Zhe Guo
The paper introduces Golden-GRPO Injection (GRIN), a three-stage self‑learning framework that uses a mixed‑policy reinforcement learning algorithm to inject knowledge into large language models. GRIN injects a golden answer to provide learning signals even when on‑policy rollouts fail on novel facts, and is evaluated on two new document‑level benchmarks—Blank and Counter—that test novel acquisition and counterfactual overwrite. Experiments show that mixed‑policy RL enables knowledge absorption beyond what supervised fine‑tuning can achieve, with GRIN outperforming SFT and other RL baselines on harder question types while matching them on basic fact recall.
By Zhibo Hou, Fan Zhao, Zhiyu An, Wan Du
CALIBURN is a new approach to large language model (LLM) unlearning that measures a model’s confidence in undesirable knowledge and uses this measure to fine‑tune unlearning gradient updates. By doing so, it offers more precise control over what is forgotten while better preserving the model’s overall utility. Experiments on benchmarks such as MUSE and WMDP show that CALIBURN outperforms existing methods in balancing knowledge removal with utility retention.
By Zhengbang Yang, Yisheng Zhong, Junyuan Hong, Zhuangdi Zhu
arXiv:2606. 15734v1 Announce Type: cross Abstract: Continual post-training enables models to absorb emerging knowledge after deployment, but repeatedly updating shared parameters can accumulate weight drift, potentially causing catastrophic forgetting and degrading general capabilities.
By Weihang Su, Jiacheng Kang, Jingyan Xu, Qingyao Ai, Jianming Long, Hanwen Zhang, Bangde Du, Xinyuan Cao, Min Zhang, Yiqun Liu
arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.
By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
The paper identifies a specific issue in supervised fine‑tuning (SFT) of large language models called factual access failure, where models can recognize correct facts under constrained tests but fail to generate them in open‑ended settings. It demonstrates that SFT can cause both genuine wrong answers and expression‑level errors such as verbosity or formatting mismatches. To mitigate this, the authors propose Recall‑Anchored Distillation (RAD), a self‑distillation method that aligns the fine‑tuned model with the base model’s soft output distribution on unlabeled out‑of‑distribution text, thereby recovering lost factual recall without needing labeled data.
By Haodong Chen, Yadong Wang, Shengtao Wen, Dong Liang, Xiang Chen