Continual Learning for Traversability Prediction with Uncertainty-Aware Adaptation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces TMLN (Trajectory-Modulatory Landscape Navigation), a method that treats continual learning as an optimal control problem on a curved loss landscape. It uses a diagonal empirical Fisher Information Matrix to approximate a local Riemannian manifold and dynamically modulates a preconditioner based on the network’s historical parameter trajectory. This trajectory‑based preconditioning is integrated into gradient updates to protect important parameter directions without adding explicit penalties, and experiments on class‑ and domain‑incremental benchmarks show a significant reduction in the loss barrier between tasks.
The paper investigates how to balance retaining past experience versus learning from new data when robot dynamics change. It introduces two metrics—change magnitude and age‑staleness AUC—to quantify when older transitions are helpful or harmful. Experiments on locomotion tasks and real‑world perturbations show that the optimal replay strategy depends on the size of the dynamics shift and the evolution of the system over time.
arXiv:2607. 17574v1 Announce Type: cross Abstract: Reinforcement-learning navigation policies for legged robots select actions reactively from current observations and short-term memory, with limited capacity to anticipate how moving obstacles will evolve in the near future.
The paper introduces ASTRIL-MPC, a language‑guided neural model predictive control framework that enables articulated tracked robots to navigate complex, contact‑rich urban environments such as stairwells and cluttered interiors. By combining a learned kinematics model that predicts short‑horizon state changes, an optimization‑based planner with multi‑objective costs, and a large language model that safely updates control weights, the system achieves up to 71% better traversal quality than non‑adaptive NMPC and 67% better than a PPO baseline, while eliminating collision impacts during descent. Real‑robot trials over four indoor obstacles confirm the method’s transferability to physical contact‑rich traversal.
The paper presents a reward‑free continual learning framework for space robots that uses latent‑state world models to adapt to severe hardware degradation. By pre‑training a model‑based agent in diverse simulations, the world model learns to predict reward structure in latent space. During deployment, the observation encoder and reward predictor are frozen while only the transition dynamics are updated via unsupervised rollouts, allowing the policy to adapt using imagined trajectories without new rewards.
World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or...