arXiv Machine Learning By Tianle Xia, Lingxiang Hu, Yiding Sun, Linfang Shang, Ming Xu, Lan Xu, Ning Zheng, Wei Xu, Jie Jiang

Distillation as Probability Transport: Routed On-Policy Distillation

Read the original on arXiv Machine Learning →

The paper introduces RouteOPD, a new on‑policy distillation method that treats teacher‑student disagreement as a probability transport problem. By decomposing disagreement into excess sources and deficit destinations, RouteOPD pairs them explicitly and optimizes pairwise log‑odds toward targets derived from a bounded teacher potential. Experiments on four teacher‑student configurations and four mathematical‑reasoning benchmarks show that RouteOPD consistently outperforms sampled reverse‑KL OPD, achieving higher routing fidelity and lower background leakage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 9

SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling

arXiv:2606. 09304v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories with dense per-token supervision from a stronger teacher, and often outperforms off-policy distillation and standard reinforcement learning.

By Haoran Xu, Hongyu Wang, Yifei Gao, Jiaze Li, Xiaofeng Zhang, Xiaosong Yuan