arXiv:2606. 02438v1 Announce Type: new Abstract: Learned heuristics have recently become a competitive alternative to traditional domain-independent heuristics for satisficing planning.
By Windy Phung, Dominik Drexler, Arnaud Lequen, Jendrik Seipp
GEM-MPC is a reinforcement learning method that blends MPPI planning with policy learning to balance exploration and exploitation in high-dimensional continuous control tasks. It trains a policy to clone the planner while also maintaining a KL-regularized policy that explores around the planner’s suggestions, thereby improving the synergy between planning and learning. The approach introduces Gated Prior Distillation, which selectively updates policies from stored planning distributions only when they offer better targets, reducing the influence of stale data without costly reanalysis. Across continuous-control benchmarks, GEM-MPC outperforms existing planning-based baselines while using lower computational budgets.
By Alvaro Serra-Gomez, Thomas Moerland
arXiv:2607. 16421v1 Announce Type: new Abstract: It has long been recognized that humans have the ability to switch between fast, reactive decision-making and slower, deliberative planning.
By Adam Labiosa, Josiah P. Hanna
arXiv:2304.10041v2 Announce Type: replace
Abstract: This work investigates formal policy synthesis for continuous-state stochastic dynamic systems subject to high-level specifications expressed in li...
By Lening Li, Zhentian Qian, Jianan Xia, Qiren Geng, Huasheng Zhang, Liang Hu, Qishuang Li, Junqiang Lou
The paper introduces an end‑to‑end model‑based reinforcement learning algorithm that synthesises policies satisfying Linear Temporal Logic (LTL) specifications in unknown environments. It synchronises a Limit‑Deterministic Büchi Automaton (LDBA) with a Bayes‑Adaptive Markov Decision Process (BAMDP) and proposes a novel Bayes‑Adaptive Monte‑Carlo Planning (BAMCP) method for approximate Bayes‑optimal strategy synthesis. Experiments on finite and infinite‑horizon tasks show improved property satisfaction and sample efficiency compared to model‑free baselines, and ablation studies confirm the advantage of the new BAMCP over classical variants, including reduced task violations in cautious RL settings.
By Jonathan Hau, Alessandro Abate
The paper introduces a system that automatically generates macro-operators—composite actions that compress recurring sequences of individual actions—by discovering causally linked action pairs in training data. It also prunes unused predicates from the symbolic state, reducing the number of predicates evaluated at each search node. These combined optimizations shorten the effective planning horizon and yield up to a 4.6× speedup, enabling the solution of long sequential tasks that baseline methods cannot solve.
By Can Emir Bora, Emre Ugur