arXiv Machine Learning By Xinmu Ge, Zizhuo Zhang, Yu Huang, Jianing Zhu, Lin Yuan, Wanli Gu, Weichang Wu, Weiran Huang, Xiaolu Zhang, Bo Han, Jun Zhou, Jiangchao Yao

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

Read the original on arXiv Machine Learning →

arXiv:2608. 11829v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 7

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

The paper investigates data efficiency and selection in On‑Policy Distillation (OPD) for large language models. It shows that 1‑shot OPD—training on a single example—consistently improves performance, especially when the example is hard, and that longer chain‑of‑thought (CoT) paths drive the gains rather than token entropy. Based on these findings, the authors propose a simple hard‑example selection strategy that, using only eight carefully chosen hard examples, matches the performance of a 17,000‑example baseline across models from 1.5B to 7B parameters.

By Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You