arXiv AI By Xiang Chen, Futao Su, Kong Wang, Jiayi Chen, TanLin Li

Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation

Read the original on arXiv AI →

The paper introduces Success-Referenced On-Policy Distillation (SR-OPD), a method that reduces teacher computation in on-policy distillation by selectively providing teacher supervision only for prompts where the student has both successful and failed rollouts. SR-OPD uses successful rollouts as references to prioritize failed rollouts that diverge significantly from the successful ones, while also considering estimated teacher-input cost. Experiments on three teacher-student pairs across six mathematical reasoning benchmarks show that SR-OPD requires only 3.46‑5.02% of the teacher-input tokens of Vanilla OPD in a one-pass setting, yet maintains comparable reasoning performance, and further validates its design choices under a 5% teacher-input budget.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 31

VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation

The paper introduces VISTA, a method that enhances on‑policy self‑distillation (OPSD) by adapting the teacher model toward the student’s distribution using outcome‑verified rollouts. VISTA keeps the standard OPSD student update but selectively adjusts the teacher only on the top‑k positions with the largest teacher‑student KL divergence, without adding new sampling or reward objectives. Experiments on AIME24, AIME25, and HMMT25 with Qwen3 models show that VISTA outperforms OPSD across all scales, improving Avg@12 by up to 2.1 points.

By Zewen Ding, Zezhong Wu, Zhou Tao, Shida Wang, Shizhuo Hou, YongXiang Hua, Haoyu Cao, Linli Xu
arXiv AI
Sep 4

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

The paper introduces Teacher-Gated On-Policy Distillation (TGOPD), a method that verifies teacher reliability at the prompt level before applying dense supervision in on-policy distillation. TGOPD uses verifier-scored teacher probes to decide whether to route a prompt to dense OPD or to a verifier-grounded alternative. Experiments on 4B and 35B models across mathematics, code, and instruction tasks show TGOPD outperforms vanilla OPD and improves teacher GPU utilization from 9.8% to 78.9% in a 4B single-domain run.

By Zhiwei Zhang, Zechen Sun, Fei Zhao, Kang Peng, Bin Liang, Huayu Deng, Yao Hu, Kam-Fai Wong, Mu Chuan