arXiv:2603. 09344v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift.
By Hongqiang Lin, Zhenghui Fu, Weihao Tang, Pengfei Wang, Yiding Sun, Qixian Huang, Dongxu Zhang
The paper introduces SUN, a reachability-aware goal-selection framework for reinforcement learning that integrates novelty and reachability using successor value functions. SUN provides theoretical guarantees, including recovery of count-based bonuses, bounds on short-horizon hitting probabilities, and rejection of unreachable goals. Empirical results show SUN consistently outperforms state-of-the-art methods across diverse environments with unreachable or hard-to-reach states, irreversible transitions, obstacles, mazes, and unbounded spaces.
By Wenyan Yang, Arsenii Mustafin, Dominik Baumann, Joni Pajarinen, Simone Parisi
arXiv:2607. 03168v1 Announce Type: cross Abstract: Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigorous theoretical foundation.
By Jialun Cao, Fernando Acero, David \v{S}i\v{s}ka, Yufei Zhang
arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.
By Zakaria Mhammedi, Alexander Rakhlin, Nneka Okolo
arXiv:2606. 10979v1 Announce Type: new Abstract: Many Markov decision processes (MDPs) in operations research have feasible actions that are state dependent and defined implicitly by various operational constraints.
By Yi Chen (Lucy), Rushuai Yang (Lucy), Qiang Chen (Lucy), Dongyan (Lucy), Huo
arXiv:2604. 00860v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a central post-training paradigm for improving the reasoning capabilities of large language models.
By Huaiyang Wang, Xiaojie Li, Deqing Wang, Haoyi Zhou, Zixuan Huang, Yaodong Yang, Jianxin Li, Yikun Ban
The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.
By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
arXiv:2605. 11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories.
By Anish Diwan, Davide Tateo, Christopher E. Mower, Haitham Bou-Ammar, Jan Peters, Oleg Arenz
arXiv:2608. 03562v1 Announce Type: new Abstract: Reinforcement learning (RL) with general utility extends classic RL by optimizing an arbitrary utility functional of the policy-induced occupancy measure, thereby enabling a broader range of applications.
By Zixuan Liu, Fangzheng Wu, Brian Summa, Zizhan Zheng
The paper introduces a two‑step deep reinforcement learning framework for Reach‑Avoid‑Stay (RAS) problems, aiming to compute the maximal robust RAS set and its control policy for general dynamic systems. First, it learns the maximal robust control‑invariant set inside the target and a policy to keep the system within it. Then it uses this invariant set as a target to compute the maximal robust reach‑avoid set, proving equivalence to the maximal robust RAS set and constructing a switching policy that guarantees task completion. Simulation results show the method achieves exact maximal RAS sets without training errors and outperforms baseline approaches in accuracy and performance.
By Gabriel Chenevert, Jingqi Li, Achyuta kannan, Sangjae Bae, Donggun Lee
arXiv:2607. 24057v1 Announce Type: new Abstract: Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy.
By Michael Girstl, Alexander Mattick, Christopher Mutschler
The paper introduces SUN, a reachability-aware goal-selection framework for reinforcement learning that jointly considers novelty and reachability. SUN uses successor value functions to identify goals that are both novel and reachable, proving properties such as recovering count-based bonuses, bounding short-horizon hitting probabilities, and rejecting unreachable goals. An adaptive goal-selection strategy and a lightweight pseudocount are proposed, and extensive benchmarks show SUN outperforming state-of-the-art methods across diverse environments.