The paper introduces Reward Stimulation Implicit Q-Learning (RSIQL), a non-hierarchical approach to improve offline goal-conditioned reinforcement learning. RSIQL adds auxiliary reward signals at intermediate states that are predicted to aid progress toward the goal, thereby reducing the delay in training supervision. Experiments on D4RL goal-reaching benchmarks and OGBench demonstrate that RSIQL outperforms baseline goal-conditioned IQL and rivals hierarchical offline methods while maintaining a simple flat policy structure.
By Jing Zhang
arXiv:2608.29061v1 Announce Type: new
Abstract: Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. Long-hori...
By Soohyun Choi, Seonvin Cho, Songnam Hong
Grounded Transitive RL (GTRL) is an offline goal‑conditioned reinforcement learning algorithm that improves upon divide‑and‑conquer value learning by grounding updates with a one‑step temporal‑difference (TD) target. By adding this TD target rather than replacing it, GTRL ensures every state‑goal pair receives an update and corrects bias from hindsight relabeling through reweighting based on reachability. The method was evaluated on nineteen OGBench tasks across stochastic, deterministic, and stitching environments, achieving the highest average success rate among compared approaches.
By Abdul Monaf Chowdhury, MD Sameer Iqbal Chowdhury, Shifat E Arman, Md Mehedi Hasan
arXiv:2602. 05459v2 Announce Type: replace Abstract: Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method.
By Jan Malte T\"opperwien, Aditya Mohan, Marius Lindauer
arXiv:2607. 20834v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets.
By Ahad Jawaid
Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets. However, scaling these methods to long-horizon tasks remains a challenge due to the curse of horizon, where value estimation errors can compound through long chains of bootstrapped Bellman backups.