Hugging Face Trending Papers

SUN: Reaching for Novelty in Reinforcement Learning

The paper introduces SUN, a reachability-aware goal-selection framework for reinforcement learning that jointly considers novelty and reachability. SUN uses successor value functions to identify goals that are both novel and reachable, proving properties such as recovering count-based bonuses, bounding short-horizon hitting probabilities, and rejecting unreachable goals. An adaptive goal-selection strategy and a lightweight pseudocount are proposed, and extensive benchmarks show SUN outperforming state-of-the-art methods across diverse environments.

arXiv AI
Sep 10

SUN: Reaching for Novelty in Reinforcement Learning

The paper introduces SUN, a reachability-aware goal-selection framework for reinforcement learning that integrates novelty and reachability using successor value functions. SUN provides theoretical guarantees, including recovery of count-based bonuses, bounds on short-horizon hitting probabilities, and rejection of unreachable goals. Empirical results show SUN consistently outperforms state-of-the-art methods across diverse environments with unreachable or hard-to-reach states, irreversible transitions, obstacles, mazes, and unbounded spaces.

By Wenyan Yang, Arsenii Mustafin, Dominik Baumann, Joni Pajarinen, Simone Parisi
arXiv Machine Learning
5d ago

Robust Successor Features

The paper introduces robust successor features, a method that extends the successor representation to handle uncertainty in both reward functions and transition kernels within linear Markov Decision Processes. It provides a theoretical bound on Generalized Policy Improvement that quantifies performance loss due to mismatched dynamics, and demonstrates the approach on grid-based benchmarks against prior methods that consider only reward or transition differences.

By Erik Nikulski, Yamen Habib, Vicen\c{c} Gomez, Anders Jonsson, Rub\'en Moreno-Bote, Javier Segovia-Aguas
arXiv AI
Jun 2

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

arXiv:2606. 00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal.

By Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno, Toshinori Kitamura, Shin Ishii, Yutaka Matsuo
arXiv AI
2d ago

Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

The paper introduces Generalized Implicit Temporal Abstraction (GITA), a method for goal-conditioned reinforcement learning that conditions a single value function on multiple temporal abstraction levels (k). By aggregating advantage-weighted supervision across various k values, GITA preserves both long-range signal and local resolution without committing to a single k. Experiments on OGBench show that GITA outperforms existing offline GCRL baselines, improving average success rates by 25 percentage points over HIQL and 7 percentage points over OTA.

By Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado
arXiv Machine Learning
Aug 28

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.

By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu
arXiv AI
2d ago

Q-Learning for Reachability in MEC-Free MDPs

The paper introduces Quasar, a model‑free Q‑learning algorithm that guarantees asymptotic convergence for reachability objectives in Markov Decision Processes that are free of non‑terminal maximal end components (MECs). Unlike prior model‑based methods, Quasar does not estimate transition probabilities, reducing memory usage from O(|S|²|A|) to O(|S||A|). Experiments on the Quantitative Verification Benchmark Set show that Quasar converges to optimal policies with far fewer samples than existing state‑of‑the‑art model‑based approaches.

By Lu-Chin Chang, Suguman Bansal
arXiv Machine Learning
Sep 11

From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs

The paper introduces Graph-Guided Quasimetric Dense Reward (G2QDR), a framework that learns a state connectivity model to predict pairwise connectivity strengths in asymmetric environments. These strengths are converted into scalar auxiliary dense rewards, offering continuous guidance across hierarchical levels. G2QDR can be integrated into any existing Goal-Conditioned Hierarchical Reinforcement Learning architecture and shows empirical performance improvements in sparse reward settings with modest computational cost.

By Shuyuan Zhang, Zihan Wang, Xiao-Wen Chang, Doina Precup
arXiv Machine Learning
1d ago

Learning Goal-Reaching Quasimetric Geometry From Finite-Time Reachability

The paper introduces ReQRL, a method for learning quasimetric geometry in goal-conditioned reinforcement learning by constraining the critic’s value gradients with finite-horizon reachability. It decouples dynamical reachability from boundary geometry, estimating both from data using state-constrained optimal control principles. Experiments on OGBench show that ReQRL matches or surpasses existing quasimetric and offline GCRL approaches.

By Daisuke Yamada, Travis Pence, Vikas Singh
arXiv Machine Learning
1d ago

When Do Intrinsic Rewards Lead to Exploration?

The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.

By Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)