arXiv:2608. 09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks.
By Tianjun Pan, Yuan Li, Hongda Wang, Linbo Jin, Mengfei Song, Lei Gao, Qiming Shi, Shaokang Fu, Jiarong Zhao, Chengyu Wang, Chengfu Huo
arXiv:2608. 13040v1 Announce Type: new Abstract: Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI.
By Guibin Zhang, Jiayang Lyu, Ran Sun, Xinlei Yu, Haoyu Zhao, Qibing Ren, Shuicheng Yan
arXiv:2607. 29468v1 Announce Type: new Abstract: Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly shape future practice.
By Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Haoyu Wu, Minghui Wu, Chenxu Zhao, Ante Wang, Guannan He, Changwei Wang
The paper introduces temporal self‑distillation for reinforcement learning with verifiable rewards (RLVR), proposing that a policy can learn from a stronger future checkpoint of itself. Two methods—Near‑Future Policy Optimization (NPO) and Near‑Future Policy Distillation (NPD)—use verified future‑self trajectories and token‑level transfer, respectively, while AutoNPO adaptively selects the optimal future checkpoint. Experiments on eight image‑text benchmarks show that near‑future teachers yield higher performance than far‑future ones, indicating that the balance between new capability and learner compatibility is key.
By Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang
Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to provide dense supervision on the student's own trajectories; however, existing methods still rely heavily on designer-specified privileged artifacts (e.
arXiv:2608. 03223v1 Announce Type: cross Abstract: Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions deserve credit.
By Ranxu Zhang, Guinan Chen, Chenshaodong, Jinghao Lin, Xiaozhou Xu, Sunzhe, Yanyong Zhang, Chao Wang
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no sup...
arXiv:2609.36642v1 Announce Type: new
Abstract: Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by lettin...
By Muyang Li, Jie Yang, Zhengyu Fang, Junchao Zhu, Zhengkun Xiao, Ruining Deng, Zhe Jiang, Shigang Chen
RetireOPD introduces a self-retiring on‑policy distillation method for agentic reinforcement learning. It first trains a skill‑conditioned teacher with environment rewards, then jointly trains a skill‑free student with RL and OPD, allowing the student to autonomously stop using the teacher when its performance aligns with the teacher’s. Experiments on Qwen2.5 models show significant gains in ALFWorld success rates and WebShop accuracy compared to RL baselines and the teacher itself.
By Yan Yu, Zhengxi Lu, Yizhou Liu, Yichen Pan, Aozhe Wang, Qipeng Chen, Hua Yang, Wenqi Zhang, Weiming Lu, Qianglong Chen, Yongliang Shen
arXiv:2607. 28076v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents.
By Binbin Zheng, Zijun Xie, Guanqun Zhao, Enlei Gong, Xing Ma, Xiaoliang Fu, Zeyu Chen
arXiv:2605.08693v3 Announce Type: replace
Abstract: Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and...
By Min Yang, Jinghua Piao, Xu Xia, Xiaochong Lan, Jiaju Chen, Yongshun Gong, Yong Li
arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.
By Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng