arXiv Machine Learning

Robust Successor Features

The paper introduces robust successor features, a method that extends the successor representation to handle uncertainty in both reward functions and transition kernels within linear Markov Decision Processes. It provides a theoretical bound on Generalized Policy Improvement that quantifies performance loss due to mismatched dynamics, and demonstrates the approach on grid-based benchmarks against prior methods that consider only reward or transition differences.

arXiv AI
Sep 10

SUN: Reaching for Novelty in Reinforcement Learning

The paper introduces SUN, a reachability-aware goal-selection framework for reinforcement learning that integrates novelty and reachability using successor value functions. SUN provides theoretical guarantees, including recovery of count-based bonuses, bounds on short-horizon hitting probabilities, and rejection of unreachable goals. Empirical results show SUN consistently outperforms state-of-the-art methods across diverse environments with unreachable or hard-to-reach states, irreversible transitions, obstacles, mazes, and unbounded spaces.

By Wenyan Yang, Arsenii Mustafin, Dominik Baumann, Joni Pajarinen, Simone Parisi
arXiv Machine Learning
Jul 1

End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions

arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.

By Zakaria Mhammedi, Alexander Rakhlin, Nneka Okolo
arXiv Machine Learning
1d ago

Towards Optimal Policy Improvement

The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.

By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
arXiv Machine Learning
Sep 3

Deep Reinforcement Learning for Reach-Avoid-Stay Problems

The paper introduces a two‑step deep reinforcement learning framework for Reach‑Avoid‑Stay (RAS) problems, aiming to compute the maximal robust RAS set and its control policy for general dynamic systems. First, it learns the maximal robust control‑invariant set inside the target and a policy to keep the system within it. Then it uses this invariant set as a target to compute the maximal robust reach‑avoid set, proving equivalence to the maximal robust RAS set and constructing a switching policy that guarantees task completion. Simulation results show the method achieves exact maximal RAS sets without training errors and outperforms baseline approaches in accuracy and performance.

By Gabriel Chenevert, Jingqi Li, Achyuta kannan, Sangjae Bae, Donggun Lee
Hugging Face Trending Papers
Sep 8

SUN: Reaching for Novelty in Reinforcement Learning

The paper introduces SUN, a reachability-aware goal-selection framework for reinforcement learning that jointly considers novelty and reachability. SUN uses successor value functions to identify goals that are both novel and reachable, proving properties such as recovering count-based bonuses, bounding short-horizon hitting probabilities, and rejecting unreachable goals. An adaptive goal-selection strategy and a lightweight pseudocount are proposed, and extensive benchmarks show SUN outperforming state-of-the-art methods across diverse environments.