arXiv AI

Reach or Solve? Deep Diving into Agentic RL Gains with Checkpoint Handoffs

The paper introduces checkpoint handoff, an evaluation protocol that separates an agent’s ability to reach useful states from its ability to complete tasks in reinforcement learning. By using one checkpoint as a reacher up to a handoff point and another as a solver from the same replayed history, the authors can measure Reach (how often states within a fixed number of actions from success are achieved) and Solve (how often the task is completed from those states). Experiments on TravelPlanner and ALFWorld show that switching the solver from supervised fine‑tuning to RL yields larger gains when RL is used as the reacher, indicating that RL more effectively finds solvable states.

arXiv AI
Sep 18

Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs

The paper introduces a new evaluation protocol called checkpoint handoff to disentangle the contributions of reaching a target state and solving the task in reinforcement learning agents. By cloning states reached by one checkpoint and handing them to another without retraining, the authors separate the REACH metric (how often a policy arrives at a state confirmed to be a fixed number of actions from success) from the SOLVE metric (how often it finishes from that identical state). Across two benchmarks and pipelines, the analysis shows that RL history benefits RL solvers more than SFT solvers, and that independent REACH and SOLVE gaps predict overall performance.

By Xuan Liu, Jingbin Qian
arXiv Computation and Language
Sep 24

Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents

The paper introduces WebMRE, an offline benchmark comprising 541 tasks and 5,293 steps extracted from WebArena trajectories, designed to provide deterministic scoring for web agents without live environments. It enables the first systematic study of how guide sentences and grounded actions reinforce each other, showing that jointly decoding a guide improves element selection accuracy and that the guide acts as a causal instruction channel. The authors fine‑tune models that outperform leading zero‑shot baselines on all offline metrics.

By Chengguang Gan, Yunhao Liang, QingHao Zhang, Shiwen Ni
arXiv AI
Sep 1

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.

By Qiancheng Zhou, Ruizhe Li
arXiv AI
Aug 26

PROOF-Gen: From Optimized Data to Better Distillation

PROOF-Gen is a method that improves distillation of tool‑calling models by recovering successful trajectories from teacher failures. It uses per‑scenario prompt optimization to generate corrective guidance that steers the teacher to a passing trajectory, then removes this guidance before training so the student learns from clean demonstrations. On τ2‑bench, PROOF-Gen recovers 93% of failed scenarios, boosting Qwen3‑4B‑Instruct‑2507’s Pass^1 from 0.132 to 0.529 and improving Gemma 4 E4B‑it by 7.2pp on BFCL v4 multi‑turn, while also raising deployed on‑device model performance by up to 5.0pp across response‑quality metrics.

By Anh Ta, Junjie Zhu, Shahin Shayandeh
arXiv AI
Sep 18

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

The paper introduces HALTER, a graph-based system that automates the reset and evaluation of long-horizon robot manipulation tasks. HALTER constructs a spatial scene graph from point clouds and vision models, uses an LLM to score rollouts, plan resets, and verify success, all without labeled success images. In experiments on a Franka arm, HALTER restores scenes in 76% of episodes, improves skill completion estimation, and reduces operator time by 72% compared to manual reset.

By Jing Jiang, Yue Yang, Xinkai Jiang, Gedas Bertasius, Daniel J. Szafir, Rudolf Lioutikov
arXiv Machine Learning
Sep 1

The Intervention Gap in Latent World Models

The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.

By Donna Vakalis
arXiv AI
Sep 2

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.

By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
arXiv AI
Sep 24

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

The paper introduces VHD-Play, a pipeline that first samples and solves a mathematical model before generating agentic reinforcement learning environments, ensuring that dynamics and evaluation are aligned from the outset. This approach yields 3,300 diverse environments at a low cost and significantly improves the performance of a large language‑model agent (Qwen3.6‑35B‑A3B) across multiple diagnostic families and external benchmarks. The study demonstrates that stateful interaction is a key factor in learning gains and that scaling the training substrate can further enhance performance.

By Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu