Offline RL with Hierarchical Action Chunking
arXiv:2607. 20834v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets.
arXiv:2606. 17377v1 Announce Type: new Abstract: We study performance-driven environment abstraction for decision-making in large Markov decision processes.
arXiv:2607. 20834v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets.
The paper introduces state abstractions that preserve the difference of Q‑functions for offline reinforcement learning, aiming to exclude irrelevant dynamics from rich state data. It proposes a dynamic generalization of the R‑learner that uses orthogonal estimation and sparse learning to estimate the Q‑function contrast, achieving faster convergence and consistency under a margin condition. Experiments on simulated and simulator‑augmented real data show variance reductions and demonstrate that the necessary information for sequential decision‑making can be smaller than that required for full state prediction.
arXiv:2606. 29806v1 Announce Type: cross Abstract: Action-values are foundational to many control algorithms such as Q-learning.
The paper introduces Generalized Implicit Temporal Abstraction (GITA), a method for goal-conditioned reinforcement learning that conditions a single value function on multiple temporal abstraction levels (k). By aggregating advantage-weighted supervision across various k values, GITA preserves both long-range signal and local resolution without committing to a single k. Experiments on OGBench show that GITA outperforms existing offline GCRL baselines, improving average success rates by 25 percentage points over HIQL and 7 percentage points over OTA.
arXiv:2602. 05031v2 Announce Type: replace Abstract: Planning with a learned model remains a key challenge in model-based reinforcement learning (RL).
arXiv:2606. 00427v1 Announce Type: new Abstract: State abstraction in reinforcement learning is usually formulated as a partition of states based on reward and transition similarity.
In value-based reinforcement learning, improving the accuracy of policy evaluation has been shown to improve downstream policy optimization performance. The widely adopted family of approximations rel...
Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets. However, scaling these methods to long-horizon tasks remains a challenge due to the curse of horizon, where value estimation errors can compound through long chains of bootstrapped Bellman backups.
arXiv:2606. 06123v1 Announce Type: new Abstract: When learning to walk, infants seem to address a coarse version of the problem first - stay upright, reach the caregiver - and refine it only when further practice at that resolution stops paying off.
arXiv:2606. 05555v1 Announce Type: new Abstract: Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge.
The paper introduces a method that integrates action abstraction into policy optimization for reinforcement learning and generative flow networks. By iteratively identifying frequently used action subsequences in high‑reward trajectories and treating them as single high‑level actions, the approach expands the action space and improves sample efficiency. Experiments on synthetic and real‑world tasks show that this technique discovers diverse high‑reward states more effectively, especially on challenging exploration problems, and yields interpretable abstract actions that reflect the underlying reward structure.
arXiv:2606. 10979v1 Announce Type: new Abstract: Many Markov decision processes (MDPs) in operations research have feasible actions that are state dependent and defined implicitly by various operational constraints.