The paper introduces a new evaluation protocol called checkpoint handoff to disentangle the contributions of reaching a target state and solving the task in reinforcement learning agents. By cloning states reached by one checkpoint and handing them to another without retraining, the authors separate the REACH metric (how often a policy arrives at a state confirmed to be a fixed number of actions from success) from the SOLVE metric (how often it finishes from that identical state). Across two benchmarks and pipelines, the analysis shows that RL history benefits RL solvers more than SFT solvers, and that independent REACH and SOLVE gaps predict overall performance.
By Xuan Liu, Jingbin Qian
arXiv:2609.01274v1 Announce Type: new
Abstract: Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but how these gains relate to inference-time decoding and sear...
By Wenhe Sun, Cunxiang Wang, Zijun Yao, Yixin Cao
arXiv:2610.07898v1 Announce Type: new
Abstract: Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt...
By Jia Liufu, Bin Hu, Linglin Jing, Terry Kong, Yuki Huang, Ashwath Aithal, Wenming Yang, Jun Yang
arXiv:2609. 40285v1 Announce Type: new Abstract: On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories.
By Yinghui He, Yapei Chang, Khushi Bhardwaj, Daniele Molinari, Tugrul Konuk, Jan Kautz, Ali Hatamizadeh
The paper introduces WebMRE, an offline benchmark comprising 541 tasks and 5,293 steps extracted from WebArena trajectories, designed to provide deterministic scoring for web agents without live environments. It enables the first systematic study of how guide sentences and grounded actions reinforce each other, showing that jointly decoding a guide improves element selection accuracy and that the guide acts as a causal instruction channel. The authors fine‑tune models that outperform leading zero‑shot baselines on all offline metrics.
By Chengguang Gan, Yunhao Liang, QingHao Zhang, Shiwen Ni
The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.
By Qiancheng Zhou, Ruizhe Li