arXiv Machine Learning

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

arXiv:2608. 11829v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning.

arXiv AI
Sep 7

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

The paper investigates data efficiency and selection in On‑Policy Distillation (OPD) for large language models. It shows that 1‑shot OPD—training on a single example—consistently improves performance, especially when the example is hard, and that longer chain‑of‑thought (CoT) paths drive the gains rather than token entropy. Based on these findings, the authors propose a simple hard‑example selection strategy that, using only eight carefully chosen hard examples, matches the performance of a 17,000‑example baseline across models from 1.5B to 7B parameters.

By Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You
arXiv AI
Aug 13

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

arXiv:2608. 12307v1 Announce Type: cross Abstract: Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods.

By Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke
arXiv Computation and Language
Sep 18

What Does Privileged Information Add to On-Policy Self-Distillation?

The paper investigates how privileged information—such as a teacher’s full solution or reasoning trace—affects on‑policy self‑distillation (OPSD) in language models. Using the AMPLE‑Math benchmark, the authors compare distillation with and without extra teacher views, finding that reference‑free distillation explains most gains for Qwen3‑1.7B, while additional references provide modest benefits, especially for polished solutions. The study also shows that the impact of privileged data depends on the student’s training regime and that altering token‑level supervision can leave student behavior largely unchanged.

By XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang, Tat-Seng Chua
arXiv AI
Sep 15

Data-free On-policy Distillation

arXiv:2609.14193v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes...

By Gengsheng Li, Mao Zheng, Mingyang Song, Jie Sun, Zeyuan Liu, Ruiqi Liu, Qiyong Zhong, Haiyun Guo, Junfeng Fang, Jinqiao Wang
arXiv AI
Aug 26

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

OPDSearch+ introduces a two‑stage distillation framework for search‑augmented reasoning that eliminates the need for task‑specific teacher fine‑tuning. In the first stage, a frozen off‑the‑shelf instruct model guides a student through live search interactions using a per‑position forward KL objective, transferring reasoning decomposition and evidence integration skills. The second stage refines this student with reinforcement learning, achieving performance surpassing RL alone and outperforming all prior 3B‑parameter baselines on seven QA benchmarks, including 13.1% improvement on HotpotQA and 8.5% on 2WikiMultihopQA.

By Qinglin Ye, Zhiyuan Gu, Jingjie Xia, Yiheng Zhang, Kaiyan Zhao, Shunchao Zheng, Yuhang Mu, Wenchao Du, Yiming Wang