The paper introduces Reward Stimulation Implicit Q-Learning (RSIQL), a non-hierarchical approach to improve offline goal-conditioned reinforcement learning. RSIQL adds auxiliary reward signals at intermediate states that are predicted to aid progress toward the goal, thereby reducing the delay in training supervision. Experiments on D4RL goal-reaching benchmarks and OGBench demonstrate that RSIQL outperforms baseline goal-conditioned IQL and rivals hierarchical offline methods while maintaining a simple flat policy structure.
By Jing Zhang
arXiv:2608.29061v1 Announce Type: new
Abstract: Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. Long-hori...
By Soohyun Choi, Seonvin Cho, Songnam Hong
Grounded Transitive RL (GTRL) is an offline goal‑conditioned reinforcement learning algorithm that improves upon divide‑and‑conquer value learning by grounding updates with a one‑step temporal‑difference (TD) target. By adding this TD target rather than replacing it, GTRL ensures every state‑goal pair receives an update and corrects bias from hindsight relabeling through reweighting based on reachability. The method was evaluated on nineteen OGBench tasks across stochastic, deterministic, and stitching environments, achieving the highest average success rate among compared approaches.
By Abdul Monaf Chowdhury, MD Sameer Iqbal Chowdhury, Shifat E Arman, Md Mehedi Hasan
arXiv:2602. 05459v2 Announce Type: replace Abstract: Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method.
By Jan Malte T\"opperwien, Aditya Mohan, Marius Lindauer
arXiv:2607. 20834v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets.
By Ahad Jawaid
Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets. However, scaling these methods to long-horizon tasks remains a challenge due to the curse of horizon, where value estimation errors can compound through long chains of bootstrapped Bellman backups.
arXiv:2606. 04188v1 Announce Type: cross Abstract: Offline goal-conditioned reinforcement learning requires both long-horizon reachability estimates and local action comparisons.
By Alexey Zemtsov, Maxim Bobrin, Alexander Nikulin, Dmitry V. Dylov, Fakhri Karray, Vladislav Kurenkov, Martin Tak\'a\v{c}, Arip Asadulaev
The paper introduces DCRL (Divide-and-Conquer RL), a method that recursively decomposes offline goal-conditioned reinforcement learning trajectories into a balanced binary tree. By training values from the leaves up to the root, DCRL avoids noisy max-based backups and reduces bootstrap depth from linear to logarithmic, thereby limiting error accumulation. Experiments on diverse goal-reaching tasks show that DCRL outperforms prior flat offline GCRL methods, achieving a higher average score on the most challenging long-horizon OGBench tasks.
By Hyeonseong Jeon, Youngwoon Lee
The paper investigates whether enhancing goal representations improves goal-conditioned reinforcement learning (GCRL) performance. By creating an exact temporal-distance goal representation in deterministic mazes and systematically degrading its geometric quality, the authors find that changes in goal representation have little effect on performance. In contrast, degrading the agent’s current state representation more than doubles failure rates, indicating that state representation is the critical bottleneck. The study further demonstrates that simple random Fourier positional encodings can significantly boost performance on challenging navigation tasks without additional map or objective modifications.
By Syed Nazmus Sakib, Abdul Monaf Chowdhury, Nafiul Haque, Shifat E Arman, Md Mehedi Hasan
arXiv:2608.30406v1 Announce Type: new
Abstract: Goal-conditioned reinforcement learning struggles with long horizons when rewards are sparse. While a planner can provide subgoals to guide a low-level...
By Olivier Serris, St\'ephane Doncieux, Olivier Sigaud
The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.
By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
The paper introduces GACA, a critic‑free reinforcement learning estimator that adapts credit assignment granularity based on a step‑level uncertainty proxy. GACA assigns higher weight to fine‑grained signals for steps with above‑average negative log‑likelihood, while relying on episode‑level signals for less uncertain steps, improving task success on ALFWorld and WebShop for 1.5B and 7B language models. The authors provide a risk decomposition, a conditional bound on action‑value variation, and an error‑projection analysis to justify the method’s effectiveness.
By Taoran Liang, Yang Liu, Shang Luo, Yingguang Yang, Rongrong Zhang, Yingzong Min, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, Congjing Ran, Bin Chong