arXiv AI

Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents

arXiv:2608. 16651v1 Announce Type: cross Abstract: Satellite agents for on-orbit navigation tasks need to predict collision risks using limited onboard observations.

arXiv Machine Learning
Aug 11

Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance

arXiv:2608. 09628v1 Announce Type: new Abstract: Collision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO) and Geosynchronous Equatorial Orbit (GEO).

By Logan Luna (Georgia Institute of Technology), Juan Ortiz Couder (Embry-Riddle Aeronautical University), Raul Alejandro Vargas-Acosta (Embry-Riddle Aeronautical University)
arXiv AI
Aug 25

Reward-Free Continual Adaptation for Resilient Space Robots

The paper presents a reward‑free continual learning framework for space robots that uses latent‑state world models to adapt to severe hardware degradation. By pre‑training a model‑based agent in diverse simulations, the world model learns to predict reward structure in latent space. During deployment, the observation encoder and reward predictor are frozen while only the transition dynamics are updated via unsupervised rollouts, allowing the policy to adapt using imagined trajectories without new rewards.

By Andrej Orsula, Miguel Olivares-Mendez, Carol Martinez
arXiv Machine Learning
Sep 16

Neural Operator Learning for Collision-Aware Trajectory Planning of Spacecraft Swarms

The paper presents a permutation‑equivariant neural operator that learns to generate collision‑free, fuel‑efficient trajectories for spacecraft swarms by mapping distributions of initial and target states, as well as obstacle states, to trajectory outputs. The operator is self‑supervised and, when paired with a batched Gauss‑Newton step, enforces exact orbital dynamics and further reduces fuel consumption. Trained on ten spacecraft, the model generalizes zero‑shot to swarms of 1,000 spacecraft and 11,000 obstacles, achieving accuracy comparable to a per‑agent optimal control solver while maintaining collision avoidance.

By Sidhdharth D. Sikka, Suyi Gao, Zehui Lu, Rongjie Lai, Shaoshuai Mou
arXiv AI
Aug 28

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

The paper introduces a Latent World Model (LWM) for robot navigation that predicts action‑conditioned latent feature compatibility instead of reconstructing future observations. By exploiting the correlation between spatial proximity and latent feature similarity, the model evaluates action consequences directly in latent space and supports counterfactual training using sampled action sequences. The learned world model can supervise policy learning from unlabeled video and further improve policies via reinforcement learning entirely within the model, eliminating the need for action annotations and additional environment interaction.

By Zengmao Wang, Wei Gao, Shuhan Shen
arXiv Machine Learning
Aug 4

Neural operator learning for collision-aware trajectory planning of spacecraft swarms

arXiv:2608. 00320v1 Announce Type: new Abstract: Autonomous spacecraft swarms must plan fuel-efficient, collision-free maneuvers in increasingly congested orbits, yet classical trajectory optimization scales poorly as pairwise safety constraints multiply with swarm size, and learning-based planners rarely transfer across swarm sizes or debris densities.

By Sidhdharth D. Sikka, Suyi Gao, Zehui Lu, Rongjie Lai, Shaoshuai Mou
arXiv Machine Learning
Sep 4

Latent Energy Action Planning with World Models

Latent Energy Action Planning (LEAP) is a new method that treats the entire action horizon as a differentiable variable and optimizes it using a frozen LeWorldModel (LeWM). LEAP couples terminal latent goal matching with a terminal‑window state energy, ensuring that the predicted terminal latent and decoder‑predicted terminal descriptor align with the goal. Using a frozen goal‑conditioned proposal, a quasi‑Newton solver, and post‑optimization projection, LEAP improves mean success from 77.5% to 94.8% across four control domains while keeping the LeWM representation frozen.

By Phu Pham, Aniket Bera