arXiv AI

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

RetireOPD introduces a self-retiring on‑policy distillation method for agentic reinforcement learning. It first trains a skill‑conditioned teacher with environment rewards, then jointly trains a skill‑free student with RL and OPD, allowing the student to autonomously stop using the teacher when its performance aligns with the teacher’s. Experiments on Qwen2.5 models show significant gains in ALFWorld success rates and WebShop accuracy compared to RL baselines and the teacher itself.

arXiv AI
1d ago

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

arXiv:2609.37898v1 Announce Type: new Abstract: Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-st...

By Youling Huang, Tiankuo Xu, Jiaji Liu, Tong Zheng, Shuo Zhou, Shaotong Qi, Junchi Yao, Shiyang Liu, Hao Xu, Pengcheng Xu, Bo Huang, Hongyi Fu, Lin Lin
arXiv AI
Jun 16

On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents

arXiv:2606. 15912v1 Announce Type: cross Abstract: Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large models whose inference cost is prohibitive in practice.

By Gengsheng Li, Mao Zheng, Mingyang Song, Ruiqi Liu, Tianyu Yang, Jie Sun, Qiyong Zhong, Haiyun Guo, Junfeng Fang, Dan Zhang, Jinqiao Wang
arXiv AI
6d ago

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

The paper introduces Privileged Self-Practice (PSP), a method that retains privileged information (PI) in the prompt rather than the loss during on‑policy self‑distillation for multi‑turn agents. PSP injects short per‑task instructions from an analyzer model when rollouts fail, sampling again with the instruction in context and training with the unchanged GRPO objective. Experiments on AppWorld and SWE‑bench Verified show PSP consistently outperforms plain GRPO, boosting task‑goal completion by up to 65% and resolved rate by up to 61% across three student models.

By Xingyu Su, Abhishek Kumar, Qing Ping, Youzhi Luo, Jonathan Buck, Zach Zhang, Subramanian Chidambaram, Vinayak Arannil
Hugging Face Trending Papers
Sep 24

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

The paper examines on‑policy self‑distillation (OPSD) for multi‑turn agents, showing that using privileged information (PI) in the loss can make agents appear confident yet underperform plain RL, sometimes worse than the untrained base model. To address this, the authors propose Privileged Self‑Practice (PSP), which keeps PI in the prompt and uses it only during sampling, not in the loss. PSP consistently outperforms plain GRPO across AppWorld and SWE‑bench Verified, improving task‑goal completion by up to 65% and resolved rate by up to 61%.

arXiv Computation and Language
3d ago

Recursive Self-Improvement via On-Policy Distillation for Reasoning

The paper introduces a recursive self-improvement framework for language models that replaces an external teacher with a frozen copy of the student, enabling dynamic co-evolution (DCE) and self-refined concise learning (SRCL). DCE allows the privileged teacher to evolve alongside the student, while SRCL trains on shorter, verified rewrites to reduce verbosity. Experiments show that the combined DCE+SRCL approach outperforms traditional on‑policy self‑distillation across multiple model sizes and math benchmarks, achieving significant accuracy gains and shorter outputs.

By Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu, Shike Mei, Hang Cui, Luke Simon, Zhouxing Shi, Hamed Firooz