arXiv:2608. 12957v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect.
By Yubo Zhang, Xinhong Ma, Zezhong Tan, Ziqiang Dong
arXiv:2609.23697v1 Announce Type: cross
Abstract: Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, howe...
By Jie Sun, Mao Zheng, Mingyang Song, Zeyuan Liu, Gengsheng Li, Houcheng Jiang, Yilin Cheng, Bichuan Feng, Yuchen Cai, Junfeng Fang, Xiang Wang
arXiv:2607. 18955v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation.
By Qiye Cai, Yichuan Ma, Linyang Li, Peiji Li, Yongkang Chen, Qipeng Guo, Yicheng Zou, Tao Gui, Xiaocheng Feng, Bing Qin
arXiv:2609.36642v1 Announce Type: new
Abstract: Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by lettin...
By Muyang Li, Jie Yang, Zhengyu Fang, Junchao Zhu, Zhengkun Xiao, Ruining Deng, Zhe Jiang, Shigang Chen
The paper investigates on‑policy distillation (OPD), showing that teacher supervision during OPD contains significant noise that grows with teacher size, yet the student policy remains largely unaffected by this noise. It finds that OPD’s gains stem mainly from suppressing low‑log‑probability tokens, a process that can be replicated without a teacher. Building on this insight, the authors propose On‑Policy Self‑Adaptation (OPSA), a supervision‑free method that uses entropy‑adaptive negative advantages to improve performance on several benchmarks, outperforming both the base model and OPD.
By Yi Ding, Ruqi Zhang
arXiv:2606. 29476v1 Announce Type: cross Abstract: Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher the same policy conditioned on privileged context.
By Zibin Meng, Kani Chen