arXiv Machine Learning

Beyond Success Rates: Trainability and Extractability for Offline GCRL

arXiv:2602. 05459v2 Announce Type: replace Abstract: Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method.

arXiv AI
Sep 4

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.

By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv AI
1d ago

Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

The paper introduces Generalized Implicit Temporal Abstraction (GITA), a method for goal-conditioned reinforcement learning that conditions a single value function on multiple temporal abstraction levels (k). By aggregating advantage-weighted supervision across various k values, GITA preserves both long-range signal and local resolution without committing to a single k. Experiments on OGBench show that GITA outperforms existing offline GCRL baselines, improving average success rates by 25 percentage points over HIQL and 7 percentage points over OTA.

By Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado
arXiv Machine Learning
Jun 25

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

arXiv:2606. 26080v1 Announce Type: new Abstract: Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estimation infeasible at scale.

By Changdae Oh, Wendi Li, Seongheon Park, Samuel Yeh, Tanwi Mallick, Sharon Li
arXiv Machine Learning
Jun 16

Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

arXiv:2606. 17043v1 Announce Type: cross Abstract: When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary outcome (success or failure), yet the actor update requires per-transition supervision.

By Tongyan Fang, Siyuan Huang, Naiyu Fang, Ganlong Zhao, Zhongjin Luo, Jianbo Liu, Xiaogang Wang, Ying Dong, Hongsheng Li
arXiv Machine Learning
Sep 18

Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation

The paper introduces Reward Stimulation Implicit Q-Learning (RSIQL), a non-hierarchical approach to improve offline goal-conditioned reinforcement learning. RSIQL adds auxiliary reward signals at intermediate states that are predicted to aid progress toward the goal, thereby reducing the delay in training supervision. Experiments on D4RL goal-reaching benchmarks and OGBench demonstrate that RSIQL outperforms baseline goal-conditioned IQL and rivals hierarchical offline methods while maintaining a simple flat policy structure.

By Jing Zhang
arXiv AI
3d ago

Beyond a single latent space: a dual-latent world model for long-horizon planning

The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.

By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
arXiv AI
Sep 2

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.

By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
arXiv AI
Jun 4

Dual Advantage Fields

arXiv:2606. 04188v1 Announce Type: cross Abstract: Offline goal-conditioned reinforcement learning requires both long-horizon reachability estimates and local action comparisons.

By Alexey Zemtsov, Maxim Bobrin, Alexander Nikulin, Dmitry V. Dylov, Fakhri Karray, Vladislav Kurenkov, Martin Tak\'a\v{c}, Arip Asadulaev