arXiv AI

Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

arXiv:2607. 09336v1 Announce Type: cross Abstract: Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost.

arXiv AI
Aug 3

RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving

arXiv:2602. 07339v2 Announce Type: replace Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment.

By Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy
arXiv AI
4d ago

HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL

HorizonFlow is a hierarchical planner for offline goal-conditioned reinforcement learning that treats the planning horizon as an output rather than a fixed input. It uses a subgoal route planner and an action-prefix controller, both employing insertion-based generation and flow matching, to jointly generate continuous plan content and its length. The method leverages the partially generated plan to guide token insertion and to steer generation toward shorter plans, achieving superior performance on Maze2D, Multi2D, and OGBench benchmarks.

By JunHyeok Oh, Zian Jang, Byung-Jun Lee
arXiv AI
Jun 10

Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning

arXiv:2606. 10613v1 Announce Type: cross Abstract: Diffusion-based Q-learning has emerged as a powerful paradigm for offline reinforcement learning, but its reliance on multi-step denoising makes both training and inference computationally expensive and brittle.

By Thanh Nguyen, Tri Ton, Hongbin Choe, Tung M. Luu, Chang D. Yoo
arXiv AI
Jun 2

Improving Diffusion Planners by Self-Supervised Action Gating with Energies

arXiv:2603. 02650v2 Announce Type: replace-cross Abstract: Diffusion planners are a strong approach for offline reinforcement learning, but they can fail when value-guided selection favours trajectories that score well yet are locally inconsistent with the environment dynamics, resulting in brittle execution.

By Yuan Lu, Dongqi Han, Yansen Wang, Dongsheng Li
arXiv Machine Learning
Sep 3

Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL

The paper introduces DCRL (Divide-and-Conquer RL), a method that recursively decomposes offline goal-conditioned reinforcement learning trajectories into a balanced binary tree. By training values from the leaves up to the root, DCRL avoids noisy max-based backups and reduces bootstrap depth from linear to logarithmic, thereby limiting error accumulation. Experiments on diverse goal-reaching tasks show that DCRL outperforms prior flat offline GCRL methods, achieving a higher average score on the most challenging long-horizon OGBench tasks.

By Hyeonseong Jeon, Youngwoon Lee
arXiv Machine Learning
Sep 21

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

GEM-MPC is a reinforcement learning method that blends MPPI planning with policy learning to balance exploration and exploitation in high-dimensional continuous control tasks. It trains a policy to clone the planner while also maintaining a KL-regularized policy that explores around the planner’s suggestions, thereby improving the synergy between planning and learning. The approach introduces Gated Prior Distillation, which selectively updates policies from stored planning distributions only when they offer better targets, reducing the influence of stale data without costly reanalysis. Across continuous-control benchmarks, GEM-MPC outperforms existing planning-based baselines while using lower computational budgets.

By Alvaro Serra-Gomez, Thomas Moerland
arXiv Machine Learning
Sep 18

Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation

The paper introduces Bidirectional Behavior Prior Distillation (B2PD), a method that uses action‑value priors to train a conditional variational autoencoder for generating high‑value behavior support. These expert behavior priors are then distilled into the online reinforcement learning agent, reducing inefficient exploration and stabilizing policy updates. Experiments on state‑ and pixel‑based tasks show that B2PD improves sample efficiency while maintaining stable learning dynamics.

By Gong Gao, Xiao Lai, Jiaji Shen, Ning Jia, Xianhui Liu, Weidong Zhao