arXiv Machine Learning

RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

The paper investigates on‑policy distillation (OPD) as a preparatory step for reinforcement learning (RL). It shows that students initialized with OPD achieve higher final RL performance than those trained directly with RL or with supervised fine‑tuning followed by RL, even when OPD offers little immediate accuracy gain. The study also finds that the choice of distillation objective (reverse‑KL vs forward‑KL) and the source of trajectories influence OPD’s effectiveness at different stages of RL training.

arXiv AI
Jun 19

Reinforcement-aware Knowledge Distillation for LLM Reasoning

arXiv:2602. 22495v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs), but the high inference cost of such models motivates distillation into smaller students.

By Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang, Dhananjay Ram, Shuo Yang, Zhuowen Tu, Wei Xia, Stefano Soatto
arXiv AI
Aug 26

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

OPDSearch+ introduces a two‑stage distillation framework for search‑augmented reasoning that eliminates the need for task‑specific teacher fine‑tuning. In the first stage, a frozen off‑the‑shelf instruct model guides a student through live search interactions using a per‑position forward KL objective, transferring reasoning decomposition and evidence integration skills. The second stage refines this student with reinforcement learning, achieving performance surpassing RL alone and outperforming all prior 3B‑parameter baselines on seven QA benchmarks, including 13.1% improvement on HotpotQA and 8.5% on 2WikiMultihopQA.

By Qinglin Ye, Zhiyuan Gu, Jingjie Xia, Yiheng Zhang, Kaiyan Zhao, Shunchao Zheng, Yuhang Mu, Wenchao Du, Yiming Wang
arXiv AI
1d ago

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

The paper introduces On‑Policy Warmup (OPW), a teacher‑guided training stage where a student agent learns from a teacher on its own interaction trajectories before switching to reinforcement learning with verifiable rewards (RLVR). OPW differs from traditional imitation by focusing on states generated by the student’s own decisions, including imperfect actions and recovery situations. The authors provide a theoretical link between on‑policy reverse‑KL distillation and trajectory‑level distribution matching, showing that, under a competent teacher and low distillation loss, OPW can lower bound initial verifier success and reduce reward‑discovery complexity, thereby accelerating RLVR performance.

By Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu
arXiv AI
Sep 4

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

The paper investigates how on-policy distillation (OPD) and reinforcement learning with verifiable rewards (RLVR) can be combined for post‑training reasoning in large language models. It shows that a two‑stage approach—first applying OPD, then RL—outperforms single‑signal methods and other joint baselines on logic and math reasoning benchmarks. The authors explain this advantage through pass@k analysis, learning dynamics, and parameter updates, concluding that OPD expands solution coverage while RL sharpens performance within that support, and that the OPD validation score is the key trigger for switching to RL.

By Boyan Li, Bingsen Chen, Chenghao Yang, Ping Nie, Chen Zhao, Xi Ye
arXiv AI
Jun 2

Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation

arXiv:2605. 09253v2 Announce Type: replace-cross Abstract: While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionately drives reasoning gains, an analogous token-level understanding of On-Policy Distillation (OPD) remains largely unexplored.

By Yuxuan Jiang, Runchao Li, Shubhashis Roy Dipta, Dawei Li, Zhao Yang