arXiv AI

RLVP: Penalize the Path, Reward the Outcome

arXiv:2607. 07435v1 Announce Type: cross Abstract: Agents acting on our behalf in the real world (e.

arXiv AI
Sep 2

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.

By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
arXiv Machine Learning
Jun 16

Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

arXiv:2606. 17043v1 Announce Type: cross Abstract: When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary outcome (success or failure), yet the actor update requires per-transition supervision.

By Tongyan Fang, Siyuan Huang, Naiyu Fang, Ganlong Zhao, Zhongjin Luo, Jianbo Liu, Xiaogang Wang, Ying Dong, Hongsheng Li
arXiv AI
Sep 2

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

The paper introduces SAGE, a framework that selectively queries a Vision‑Language Model (VLM) teacher only when the learner is uncertain, using the teacher’s suggestions to guide training and distill them into a lightweight reinforcement learning policy. SAGE weights teacher actions by environment‑derived advantages, allowing the policy to improve beyond the imperfect VLM. Experiments on sparse‑reward visual reasoning and navigation tasks show that the learned policies can act without VLM guidance at evaluation, reduce VLM usage during training, and sometimes outperform the teacher itself.

By Matteo Merler, Giovanni Bonetta, Davide Zago, Rossella Cancelliere, Bernardo Magnini
arXiv AI
Jul 7

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

arXiv:2607. 03702v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajectories: full retries incur high interaction costs, while experience retrieval tends to dilute critical experience signals.

By Weiyang Guo, Zesheng Shi, Longhui Zhang, Zeen Zhu, Min Zhang, Jing Li
arXiv AI
Sep 21

Prompt-Driven Exploration: Language as an Exploration Space for VLA Reinforcement Learning

arXiv:2607.08837v4 Announce Type: replace-cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject...

By Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit Agrawal, Zhang-Wei Hong
arXiv AI
Sep 25

Reinforcement Learning with Verifiable Rewards for Small Search Agents

The paper introduces Reinforcement Learning with Verifiable Rewards (RLVR) applied to small search agents, specifically training a Qwen3.5-0.8B model with Group Relative Policy Optimization and an interleaved Wikipedia-search tool on the MuSiQue dataset. Experiments varying reward shapes across three seeds show that RLVR can achieve a 3.8‑fold improvement over an untrained baseline, with the best run reaching a 0.352 average exact match. The study finds that the sparse exact‑match reward, standard in larger models, performs poorly for small models, indicating that reward design must be tailored rather than scaled down from large‑model recipes.

By Gaurisankar Jayadas, Aske Plaat, \'Alvaro Serra-G\'omez, Sandheep P
arXiv AI
3d ago

Trust the Critic More

The paper introduces Actor‑Critic with Action Chunking (AC2), a method that assigns credit to short action chunks instead of entire trajectories, enabling policy updates without waiting for terminal rewards. AC2 employs local readiness, reference solutions, and 10k‑token chunks to make critic‑based credit assignment reliable. Experiments on Qwen3‑4B with FineProofs‑RL show AC2 surpasses GRPO’s peak validation score while using 2.5× fewer decoding FLOPs and fewer training steps.

By Kaiyue Wen, Luke Bailey, Arvind Mahankali, Tengyu Ma