arXiv Machine Learning By Siddharth Aphale, Ayushman Singh

SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2607. 26417v1 Announce Type: new Abstract: Sparse-reward reinforcement learning often fails because rollouts from the unassisted evaluation start rarely reach later task stages.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
4d ago

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

arXiv:2609.37898v1 Announce Type: new Abstract: Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-st...

By Youling Huang, Tiankuo Xu, Jiaji Liu, Tong Zheng, Shuo Zhou, Shaotong Qi, Junchi Yao, Shiyang Liu, Hao Xu, Pengcheng Xu, Bo Huang, Hongyi Fu, Lin Lin
arXiv AI
Sep 2

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.

By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
arXiv AI
Sep 15

Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training

The paper introduces an exploration-guided prompt scaffolding framework for multimodal large language models, dynamically adjusting the prompt distribution during reinforcement learning post-training. It uses an Exploration Potential Score (EPS) derived from KL-regularized policy improvement to assess prompt utility without extra overhead, and a teacher model rewrites low-utility prompts to preserve intent while improving informativeness. Experiments on Geo3K, MMK12, MathVision, and MMMU-Pro show consistent performance gains, up to 9.7% in-domain and over 11% on out-of-distribution benchmarks.

By Yuanhao Yue, Qianli Ma, Chengyu Wang, Haoting Wang, Lei Shen, Jun Huang
arXiv AI
Sep 18

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

The paper introduces HALTER, a graph-based system that automates the reset and evaluation of long-horizon robot manipulation tasks. HALTER constructs a spatial scene graph from point clouds and vision models, uses an LLM to score rollouts, plan resets, and verify success, all without labeled success images. In experiments on a Franka arm, HALTER restores scenes in 76% of episodes, improves skill completion estimation, and reduces operator time by 72% compared to manual reset.

By Jing Jiang, Yue Yang, Xinkai Jiang, Gedas Bertasius, Daniel J. Szafir, Rudolf Lioutikov