arXiv AI By Qitai Tan, Zefang Zong, Yang Li, Peng Chen

ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents

Read the original on arXiv AI →

arXiv:2606. 27814v1 Announce Type: new Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

The paper introduces On‑Policy Warmup (OPW), a teacher‑guided training stage where a student agent learns from a teacher on its own interaction trajectories before switching to reinforcement learning with verifiable rewards (RLVR). OPW differs from traditional imitation by focusing on states generated by the student’s own decisions, including imperfect actions and recovery situations. The authors provide a theoretical link between on‑policy reverse‑KL distillation and trajectory‑level distribution matching, showing that, under a competent teacher and low distillation loss, OPW can lower bound initial verifier success and reduce reward‑discovery complexity, thereby accelerating RLVR performance.

By Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu
arXiv AI
4d ago

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

arXiv:2609.37898v1 Announce Type: new Abstract: Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-st...

By Youling Huang, Tiankuo Xu, Jiaji Liu, Tong Zheng, Shuo Zhou, Shaotong Qi, Junchi Yao, Shiyang Liu, Hao Xu, Pengcheng Xu, Bo Huang, Hongyi Fu, Lin Lin
arXiv AI
Aug 26

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

OPDSearch+ introduces a two‑stage distillation framework for search‑augmented reasoning that eliminates the need for task‑specific teacher fine‑tuning. In the first stage, a frozen off‑the‑shelf instruct model guides a student through live search interactions using a per‑position forward KL objective, transferring reasoning decomposition and evidence integration skills. The second stage refines this student with reinforcement learning, achieving performance surpassing RL alone and outperforming all prior 3B‑parameter baselines on seven QA benchmarks, including 13.1% improvement on HotpotQA and 8.5% on 2WikiMultihopQA.

By Qinglin Ye, Zhiyuan Gu, Jingjie Xia, Yiheng Zhang, Kaiyan Zhao, Shunchao Zheng, Yuhang Mu, Wenchao Du, Yiming Wang