arXiv AI By Syed Nazmus Sakib, Abdul Monaf Chowdhury, Nafiul Haque, Shifat E Arman, Md Mehedi Hasan

Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?

Read the original on arXiv AI →

The paper investigates whether enhancing goal representations improves goal-conditioned reinforcement learning (GCRL) performance. By creating an exact temporal-distance goal representation in deterministic mazes and systematically degrading its geometric quality, the authors find that changes in goal representation have little effect on performance. In contrast, degrading the agent’s current state representation more than doubles failure rates, indicating that state representation is the critical bottleneck. The study further demonstrates that simple random Fourier positional encodings can significantly boost performance on challenging navigation tasks without additional map or objective modifications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL

HorizonFlow is a hierarchical planner for offline goal-conditioned reinforcement learning that treats the planning horizon as an output rather than a fixed input. It uses a subgoal route planner and an action-prefix controller, both employing insertion-based generation and flow matching, to jointly generate continuous plan content and its length. The method leverages the partially generated plan to guide token insertion and to steer generation toward shorter plans, achieving superior performance on Maze2D, Multi2D, and OGBench benchmarks.

By JunHyeok Oh, Zian Jang, Byung-Jun Lee
arXiv Machine Learning
Aug 17

CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving

arXiv:2608. 14332v1 Announce Type: cross Abstract: Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixed objective gives no order in which to learn them.

By Anisa Saleem, Duksu Kim
arXiv Statistics ML
1d ago

Learning to Plan from Random Exploration

arXiv:2609.38383v1 Announce Type: cross Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without po...

By Deqian Kong, Guangyan Sun, Sheng Cheng, Sirui Xie, Bo Pang, Jianwen Xie, Tony Geng, Caiwen Ding, Ying Nian Wu