arXiv AI By Irene Brugnara, Alessandro Valentini, Andrea Micheli

Exploiting Symbolic Heuristics for the Synthesis of Domain-Specific Temporal Planning Guidance using Reinforcement Learning

Read the original on arXiv AI →

arXiv:2505. 13372v2 Announce Type: replace Abstract: Recent work investigated the use of Reinforcement Learning (RL) for the synthesis of heuristic guidance to improve the performance of temporal planners when a domain is fixed and a set of training problems (not plans) is given.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 21

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

GEM-MPC is a reinforcement learning method that blends MPPI planning with policy learning to balance exploration and exploitation in high-dimensional continuous control tasks. It trains a policy to clone the planner while also maintaining a KL-regularized policy that explores around the planner’s suggestions, thereby improving the synergy between planning and learning. The approach introduces Gated Prior Distillation, which selectively updates policies from stored planning distributions only when they offer better targets, reducing the influence of stale data without costly reanalysis. Across continuous-control benchmarks, GEM-MPC outperforms existing planning-based baselines while using lower computational budgets.

By Alvaro Serra-Gomez, Thomas Moerland
arXiv Machine Learning
Sep 21

Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications

The paper introduces an end‑to‑end model‑based reinforcement learning algorithm that synthesises policies satisfying Linear Temporal Logic (LTL) specifications in unknown environments. It synchronises a Limit‑Deterministic Büchi Automaton (LDBA) with a Bayes‑Adaptive Markov Decision Process (BAMDP) and proposes a novel Bayes‑Adaptive Monte‑Carlo Planning (BAMCP) method for approximate Bayes‑optimal strategy synthesis. Experiments on finite and infinite‑horizon tasks show improved property satisfaction and sample efficiency compared to model‑free baselines, and ablation studies confirm the advantage of the new BAMCP over classical variants, including reduced task violations in cautious RL settings.

By Jonathan Hau, Alessandro Abate
arXiv AI
Aug 26

Macro-Operator Generation and Predicate Selection for TAMP Operator Learning

The paper introduces a system that automatically generates macro-operators—composite actions that compress recurring sequences of individual actions—by discovering causally linked action pairs in training data. It also prunes unused predicates from the symbolic state, reducing the number of predicates evaluated at each search node. These combined optimizations shorten the effective planning horizon and yield up to a 4.6× speedup, enabling the solution of long sequential tasks that baseline methods cannot solve.

By Can Emir Bora, Emre Ugur