arXiv AI By Jiaxuan Wang, Jiafei Lyu, Yuchen Cai, Siye Wu, Pengyuan Wang, Jiashun Liu, Xiang Cheng, Kai Yang, Yangkun Chen, Saiyong Yang, Lan-Zhe Guo

Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation

Read the original on arXiv AI →

The paper investigates how the quantity, source, and selection of prompts influence transfer in on‑policy distillation (OPD) between teacher and student models. It shows that a small set of well‑chosen prompts can achieve performance comparable to large prompt pools, but the effectiveness of prompts depends on the specific teacher‑student pair and target task. The study also finds that prompt utility is relational rather than intrinsic, and that targeted prompt selection does not consistently outperform random sampling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
2d ago

No Task Vector Is an Island: A Comprehensive Study on the Composability of Task Vectors from On-Policy Distillation

arXiv:2609.39405v1 Announce Type: new Abstract: Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced b...

By Jingang Zhou, Feiyu Han, Han Zhu, Yuyi Zhou, Ruiyang Zhang, Jian Xu, Sirui Gao, Qingpei Guo, Xu-Yao Zhang
arXiv Computation and Language
Sep 23

Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation

The study investigates how the number of prompts and the strategy of refreshing rollout responses affect on‑policy distillation (OPD). Using a 3×3 experiment with 14,080 trajectories and 110 optimizer updates, the authors find that with ten policy snapshots, eight prompts achieve 24.09% accuracy—nearly matching the 24.51% obtained with 14,080 distinct prompts. However, when responses are frozen at the initial policy, increasing prompt breadth actually reduces accuracy, whereas per‑update refresh raises it, producing a 4.07‑point interaction effect. Comparisons with two teacher models show that periodic models excel in short‑budget accuracy and answer completion, but frozen‑response models surpass them in overall accuracy at a 32K output limit, using 1.7–1.8× more response tokens. whyItMatters":"The findings demonstrate that prompt efficiency in OPD is contingent on both the refresh strategy and the inference budget, informing how to design more effective distillation pipelines."

By Lingxiang Hu, Tianle Xia, Ming Xu, Yiding Sun, Linfang Shang
arXiv Machine Learning
2d ago

Activation-Conditioned Self-Distillation

arXiv:2609.38342v1 Announce Type: new Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....

By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu