arXiv AI By Xuan Liu, Jingbin Qian

Reach or Solve? Deep Diving into Agentic RL Gains with Checkpoint Handoffs

Read the original on arXiv AI →

The paper introduces checkpoint handoff, an evaluation protocol that separates an agent’s ability to reach useful states from its ability to complete tasks in reinforcement learning. By using one checkpoint as a reacher up to a handoff point and another as a solver from the same replayed history, the authors can measure Reach (how often states within a fixed number of actions from success are achieved) and Solve (how often the task is completed from those states). Experiments on TravelPlanner and ALFWorld show that switching the solver from supervised fine‑tuning to RL yields larger gains when RL is used as the reacher, indicating that RL more effectively finds solvable states.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs

The paper introduces a new evaluation protocol called checkpoint handoff to disentangle the contributions of reaching a target state and solving the task in reinforcement learning agents. By cloning states reached by one checkpoint and handing them to another without retraining, the authors separate the REACH metric (how often a policy arrives at a state confirmed to be a fixed number of actions from success) from the SOLVE metric (how often it finishes from that identical state). Across two benchmarks and pipelines, the analysis shows that RL history benefits RL solvers more than SFT solvers, and that independent REACH and SOLVE gaps predict overall performance.

By Xuan Liu, Jingbin Qian
arXiv Computation and Language
Sep 24

Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents

The paper introduces WebMRE, an offline benchmark comprising 541 tasks and 5,293 steps extracted from WebArena trajectories, designed to provide deterministic scoring for web agents without live environments. It enables the first systematic study of how guide sentences and grounded actions reinforce each other, showing that jointly decoding a guide improves element selection accuracy and that the guide acts as a causal instruction channel. The authors fine‑tune models that outperform leading zero‑shot baselines on all offline metrics.

By Chengguang Gan, Yunhao Liang, QingHao Zhang, Shiwen Ni
arXiv AI
Sep 1

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.

By Qiancheng Zhou, Ruizhe Li