arXiv:2607. 07050v1 Announce Type: cross Abstract: Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly.
By Jiabin Shen, Guang Chen, Chengjun Mao
arXiv:2607. 07050v2 Announce Type: replace-cross Abstract: Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly.
By Jiabin Shen, Guang Chen, Chengjun Mao
arXiv:2608. 03632v1 Announce Type: new Abstract: On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals.
By Yinuo Jiang, Yongjie Ye, Zhou Tao, Xiang Zhuang, Qiang Zhang, Huajun Chen, Tiankai Li
arXiv:2607. 02502v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access.
By Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable.
arXiv:2607. 18293v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) teaches large language models new skills through a teacher that shares the student's backbone and supervises its own rollouts.
By Yingzi Ma, Zichen Zhu, Ming Jiang, Chaowei Xiao