arXiv Machine Learning
Jul 22

H$^2$SD: Hybrid Hindsight Self-Distillation

arXiv:2607. 18955v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation.

By Qiye Cai, Yichuan Ma, Linyang Li, Peiji Li, Yongkang Chen, Qipeng Guo, Yicheng Zou, Tao Gui, Xiaocheng Feng, Bing Qin
arXiv AI
Sep 18

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

RetireOPD introduces a self-retiring on‑policy distillation method for agentic reinforcement learning. It first trains a skill‑conditioned teacher with environment rewards, then jointly trains a skill‑free student with RL and OPD, allowing the student to autonomously stop using the teacher when its performance aligns with the teacher’s. Experiments on Qwen2.5 models show significant gains in ALFWorld success rates and WebShop accuracy compared to RL baselines and the teacher itself.

By Yan Yu, Zhengxi Lu, Yizhou Liu, Yichen Pan, Aozhe Wang, Qipeng Chen, Hua Yang, Wenqi Zhang, Weiming Lu, Qianglong Chen, Yongliang Shen
arXiv AI
6d ago

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

The paper introduces Privileged Self-Practice (PSP), a method that retains privileged information (PI) in the prompt rather than the loss during on‑policy self‑distillation for multi‑turn agents. PSP injects short per‑task instructions from an analyzer model when rollouts fail, sampling again with the instruction in context and training with the unchanged GRPO objective. Experiments on AppWorld and SWE‑bench Verified show PSP consistently outperforms plain GRPO, boosting task‑goal completion by up to 65% and resolved rate by up to 61% across three student models.

By Xingyu Su, Abhishek Kumar, Qing Ping, Youzhi Luo, Jonathan Buck, Zach Zhang, Subramanian Chidambaram, Vinayak Arannil
arXiv AI
Aug 6

Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation

arXiv:2608. 04794v1 Announce Type: new Abstract: Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI) about the answer such as a reference solution, supplies dense per-token supervision to a student that never sees it.

By Sarthak Harne, Chinmay Karkar, Yash Pandya, Ahmed Awadallah, Akshay Nambi