arXiv Machine Learning

Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers

The paper introduces Reachability Analysis-Informed Reinforcement Learning (RARL) for designing deterministic multi‑impulse interplanetary transfers. RARL uses local first‑order reachability maps to bound velocity perturbations and selects intermediate waypoints, which are then translated into maneuvers via Lambert reconstruction and a terminal two‑impulse solution. Numerical experiments on an Earth‑Mars benchmark show that RARL achieves a mean maneuver cost only 1.72% above a validated convex programming reference and can be trained once to handle a wide range of departure states, achieving 100% feasibility on 10,000 held‑out Monte Carlo departures.

arXiv Machine Learning
Aug 11

Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance

arXiv:2608. 09628v1 Announce Type: new Abstract: Collision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO) and Geosynchronous Equatorial Orbit (GEO).

By Logan Luna (Georgia Institute of Technology), Juan Ortiz Couder (Embry-Riddle Aeronautical University), Raul Alejandro Vargas-Acosta (Embry-Riddle Aeronautical University)
arXiv AI
Sep 16

Learning-Guided Planning in Large Dynamic Action Spaces: Budgeted Tree Search for One-to-Many Mobile Charging

The paper introduces LP‑BTS, a learning‑guided planning framework for mobile charging in large, dynamic action spaces. It uses a graph proposal policy to narrow candidate stops, a value critic to evaluate leaf nodes, and edge‑budgeted PUCT to compare short simulated futures before action selection. Experiments on a 30‑scenario battery‑life benchmark show LP‑BTS achieving the highest survival and alive‑AUC, outperforming domain‑engineered baselines and heuristic policies.

By Liang-Ching Tao, Pi-Chung Wang
arXiv Machine Learning
Aug 11

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

arXiv:2608. 07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode, structural parameter, logarithmic normalization, prior information, and planning assumptions differ.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
arXiv Machine Learning
Aug 17

CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving

arXiv:2608. 14332v1 Announce Type: cross Abstract: Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixed objective gives no order in which to learn them.

By Anisa Saleem, Duksu Kim