arXiv AI By Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado

Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

Read the original on arXiv AI →

The paper introduces Generalized Implicit Temporal Abstraction (GITA), a method for goal-conditioned reinforcement learning that conditions a single value function on multiple temporal abstraction levels (k). By aggregating advantage-weighted supervision across various k values, GITA preserves both long-range signal and local resolution without committing to a single k. Experiments on OGBench show that GITA outperforms existing offline GCRL baselines, improving average success rates by 25 percentage points over HIQL and 7 percentage points over OTA.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 18

Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation

The paper introduces Reward Stimulation Implicit Q-Learning (RSIQL), a non-hierarchical approach to improve offline goal-conditioned reinforcement learning. RSIQL adds auxiliary reward signals at intermediate states that are predicted to aid progress toward the goal, thereby reducing the delay in training supervision. Experiments on D4RL goal-reaching benchmarks and OGBench demonstrate that RSIQL outperforms baseline goal-conditioned IQL and rivals hierarchical offline methods while maintaining a simple flat policy structure.

By Jing Zhang
arXiv Machine Learning
3d ago

GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences

Grounded Transitive RL (GTRL) is an offline goal‑conditioned reinforcement learning algorithm that improves upon divide‑and‑conquer value learning by grounding updates with a one‑step temporal‑difference (TD) target. By adding this TD target rather than replacing it, GTRL ensures every state‑goal pair receives an update and corrects bias from hindsight relabeling through reweighting based on reachability. The method was evaluated on nineteen OGBench tasks across stochastic, deterministic, and stitching environments, achieving the highest average success rate among compared approaches.

By Abdul Monaf Chowdhury, MD Sameer Iqbal Chowdhury, Shifat E Arman, Md Mehedi Hasan
Hugging Face Trending Papers
Jul 23

Offline RL with Hierarchical Action Chunking

Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets. However, scaling these methods to long-horizon tasks remains a challenge due to the curse of horizon, where value estimation errors can compound through long chains of bootstrapped Bellman backups.