Hugging Face Trending Papers

Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning

This paper presents a distribution-agnostic robust trajectory-optimization framework based on chance-constrained reinforcement learning. The uncertainty is represented here through initial conditions and process noise, with the only requirement being that it can be sampled.

arXiv Machine Learning
1d ago

Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers

The paper introduces Reachability Analysis-Informed Reinforcement Learning (RARL) for designing deterministic multi‑impulse interplanetary transfers. RARL uses local first‑order reachability maps to bound velocity perturbations and selects intermediate waypoints, which are then translated into maneuvers via Lambert reconstruction and a terminal two‑impulse solution. Numerical experiments on an Earth‑Mars benchmark show that RARL achieves a mean maneuver cost only 1.72% above a validated convex programming reference and can be trained once to handle a wide range of departure states, achieving 100% feasibility on 10,000 held‑out Monte Carlo departures.

By Yashdeep Chaudhary, Roberto Armellin, Harry Holt
arXiv Machine Learning
5d ago

Learning Chance-Constrained MDPs with Bellman Distributional Certificates

The paper introduces a new approach to learning chance-constrained Markov decision processes (CCMDPs) using a Bellman distributional certificate. It provides both model-based and model-free algorithms with theoretical guarantees, including matching upper and lower bounds for tabular discounted CCMDPs with bounded successor support. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage benchmark demonstrate the safety and effectiveness of the proposed methods.

By Chenbei Lu, Hongyu Yi
arXiv Machine Learning
Sep 3

Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning

The paper introduces Exchange Policy Optimization (EPO), a framework for semi‑infinite safe reinforcement learning that handles infinitely many constraints by iteratively solving finite subproblems. EPO expands or deletes constraints based on tolerance violations and Lagrange multipliers, maintaining computational tractability while converging to an optimal policy with bounded safety violations. The authors prove finite convergence, provide iteration bounds, and quantify the suboptimality gap under mild assumptions.

By Jiaming Zhang, Yujie Yang, Haoning Wang, Liping Zhang, Shengbo Eben Li
arXiv Machine Learning
Sep 18

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

The paper introduces a model-based bootstrap framework for uncertainty quantification in offline policy evaluation (OPE) within finite-horizon, time-inhomogeneous Markov decision processes. Unlike traditional bootstrap methods that resample entire episodes, this approach regenerates trajectories from an estimated MDP, enabling use of diverse offline data formats such as complete trajectories, transition-level observations, and trajectory fragments. The authors prove bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation, and demonstrate through simulations that the method yields tighter confidence intervals and more accurate variance estimates compared to existing techniques.

By Weiwei Wang, Yuqiang Li, Xianyi Wu, Bingyi Jing
arXiv Machine Learning
Jul 7

Optimality-Informed Neural Networks for Lunar Landing Trajectory Optimization

arXiv:2607. 02741v1 Announce Type: cross Abstract: This paper develops an Optimality-Informed Neural Network (OINN) approach for the energy-optimal, free-final-time powered descent of a lunar lander from any initial position, velocity, and mass within a bounded operating envelope to a fixed landing site with zero terminal velocity.

By Zhenbo Wang