arXiv AI

Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation

The paper investigates how the quantity, source, and selection of prompts influence transfer in on‑policy distillation (OPD) between teacher and student models. It shows that a small set of well‑chosen prompts can achieve performance comparable to large prompt pools, but the effectiveness of prompts depends on the specific teacher‑student pair and target task. The study also finds that prompt utility is relational rather than intrinsic, and that targeted prompt selection does not consistently outperform random sampling.

arXiv Machine Learning
2d ago

No Task Vector Is an Island: A Comprehensive Study on the Composability of Task Vectors from On-Policy Distillation

arXiv:2609.39405v1 Announce Type: new Abstract: Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced b...

By Jingang Zhou, Feiyu Han, Han Zhu, Yuyi Zhou, Ruiyang Zhang, Jian Xu, Sirui Gao, Qingpei Guo, Xu-Yao Zhang
arXiv Computation and Language
Sep 23

Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation

The study investigates how the number of prompts and the strategy of refreshing rollout responses affect on‑policy distillation (OPD). Using a 3×3 experiment with 14,080 trajectories and 110 optimizer updates, the authors find that with ten policy snapshots, eight prompts achieve 24.09% accuracy—nearly matching the 24.51% obtained with 14,080 distinct prompts. However, when responses are frozen at the initial policy, increasing prompt breadth actually reduces accuracy, whereas per‑update refresh raises it, producing a 4.07‑point interaction effect. Comparisons with two teacher models show that periodic models excel in short‑budget accuracy and answer completion, but frozen‑response models surpass them in overall accuracy at a 32K output limit, using 1.7–1.8× more response tokens. whyItMatters":"The findings demonstrate that prompt efficiency in OPD is contingent on both the refresh strategy and the inference budget, informing how to design more effective distillation pipelines."

By Lingxiang Hu, Tianle Xia, Ming Xu, Yiding Sun, Linfang Shang
arXiv Machine Learning
2d ago

Activation-Conditioned Self-Distillation

arXiv:2609.38342v1 Announce Type: new Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....

By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv Machine Learning
Sep 24

RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

The paper investigates on‑policy distillation (OPD) as a preparatory step for reinforcement learning (RL). It shows that students initialized with OPD achieve higher final RL performance than those trained directly with RL or with supervised fine‑tuning followed by RL, even when OPD offers little immediate accuracy gain. The study also finds that the choice of distillation objective (reverse‑KL vs forward‑KL) and the source of trajectories influence OPD’s effectiveness at different stages of RL training.

By Shuai Dong, Yongfu Zhu, Yuqi Xu, Weichu Xie, Liuwenpu, Ziyue Wang, Kaiwen Tuo, Congcong Wang, Siyuan Wang, Wenqi Shao, Shuai Yang, Ji Zhao, Caoyuan Ma, Wenzheng Chang, Taiqiang Wu, Xinlei Yu, Hongrui Wu, Xiaoxuan He, Fangke Chen, Dianyi Wang, Kanghui Tian, Sirry Chen, Xingyu Liu, Xiangnan Wu, Jiawei Guo, Haowen Hou, LingHan Chen, Zhongyu Wei, Jiaqi Wang
arXiv AI
Aug 24

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

The paper introduces Preference‑Based Self‑Distillation (PBSD), a new on‑policy self‑distillation method that replaces traditional KL matching with a reward‑regularized objective. PBSD derives a reward‑reweighted teacher distribution, optimizing preference gaps between teacher and student samples while keeping on‑policy sampling. Experiments on mathematical reasoning and tool‑use tasks show PBSD achieves stronger average performance, improved training stability, and maintains token efficiency compared to prior self‑distillation baselines.

By Xin Yu, Liuchen Liao, Yiwen Zhang, Yingchen Yu, Lingzhou Xue, Qinzhen Guo
arXiv AI
Aug 18

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

arXiv:2608. 16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories.

By Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu