The paper introduces TMLN (Trajectory-Modulatory Landscape Navigation), a method that treats continual learning as an optimal control problem on a curved loss landscape. It uses a diagonal empirical Fisher Information Matrix to approximate a local Riemannian manifold and dynamically modulates a preconditioner based on the network’s historical parameter trajectory. This trajectory‑based preconditioning is integrated into gradient updates to protect important parameter directions without adding explicit penalties, and experiments on class‑ and domain‑incremental benchmarks show a significant reduction in the loss barrier between tasks.
By Isabelle Aguilar, Zayn Andre Zainal, Luis Fernando Herbozo Contreras, Zhaojing Huang, Omid Kavehei
The paper investigates how to balance retaining past experience versus learning from new data when robot dynamics change. It introduces two metrics—change magnitude and age‑staleness AUC—to quantify when older transitions are helpful or harmful. Experiments on locomotion tasks and real‑world perturbations show that the optimal replay strategy depends on the size of the dynamics shift and the evolution of the system over time.
By Everest Yang, Skye Thompson, George D. Konidaris
arXiv:2607. 17574v1 Announce Type: cross Abstract: Reinforcement-learning navigation policies for legged robots select actions reactively from current observations and short-term memory, with limited capacity to anticipate how moving obstacles will evolve in the near future.
By Yancheng Zhu, Wanli Ma, Chen Han, Irvin Haozhe Zhan, Bingfeng Qin, Yixin Xu
The paper introduces ASTRIL-MPC, a language‑guided neural model predictive control framework that enables articulated tracked robots to navigate complex, contact‑rich urban environments such as stairwells and cluttered interiors. By combining a learned kinematics model that predicts short‑horizon state changes, an optimization‑based planner with multi‑objective costs, and a large language model that safely updates control weights, the system achieves up to 71% better traversal quality than non‑adaptive NMPC and 67% better than a PPO baseline, while eliminating collision impacts during descent. Real‑robot trials over four indoor obstacles confirm the method’s transferability to physical contact‑rich traversal.
By Zhenfeng Gan, Yanbo Chen, Lirong Che, Yongyi Ma, Rongkai Zhu, Xueqian Wang
The paper presents a reward‑free continual learning framework for space robots that uses latent‑state world models to adapt to severe hardware degradation. By pre‑training a model‑based agent in diverse simulations, the world model learns to predict reward structure in latent space. During deployment, the observation encoder and reward predictor are frozen while only the transition dynamics are updated via unsupervised rollouts, allowing the policy to adapt using imagined trajectories without new rewards.
By Andrej Orsula, Miguel Olivares-Mendez, Carol Martinez
World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or...
arXiv:2607. 13579v1 Announce Type: cross Abstract: Enabling quadrupedal robots to traverse complex terrains-from rugged outdoor environments to urban landscapes-requires seamless integration of multiple motor skills, smooth transitions between gaits, and high-speed perceptive locomotion using only onboard sensors.
By Jun-Gill Kang, Jaehyun Park, Tae-Gyu Song, Joon-Ha Kim, Seungwoo Hong, Hae-Won Park
arXiv:2609.37476v1 Announce Type: cross
Abstract: Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environme...
By Jiaming Wang, Duc Thang Nguyen, Jizhuo Chen, Volodymyr Shcherbyna, Diwen Liu, Zhengcheng Shen, Harold Soh
The paper introduces a compositional continual learning benchmark for world models in robot manipulation, designed to isolate knowledge reuse from learning speed and capacity. Tasks are curated to combine previously seen action and perception components, allowing analysis of how different modalities affect reuse. Experiments show that modular world models better balance reuse and forgetting than conventional methods, yet none fully solve the challenge, highlighting the need for models explicitly built to reuse knowledge without forgetting.
By Haoyu Zhou, Joe Watson, Anson Lei, Ingmar Posner
arXiv:2608. 04334v1 Announce Type: cross Abstract: Contemporary model-free reinforcement learning algorithms can achieve very high performance, but have low sample efficiency and are not robust to changes in the environment.
By R. Blake Lawlor, Daniel S. Brown
The paper introduces a Latent World Model (LWM) for robot navigation that predicts action‑conditioned latent feature compatibility instead of reconstructing future observations. By exploiting the correlation between spatial proximity and latent feature similarity, the model evaluates action consequences directly in latent space and supports counterfactual training using sampled action sequences. The learned world model can supervise policy learning from unlabeled video and further improve policies via reinforcement learning entirely within the model, eliminating the need for action annotations and additional environment interaction.
By Zengmao Wang, Wei Gao, Shuhan Shen
arXiv:2601. 15995v2 Announce Type: replace-cross Abstract: Parkour tasks for quadrupeds have emerged as a promising benchmark for agile locomotion.
By Liang Wang, Kanzhong Yao, Yang Liu, Weikai Qin, Jun Wu, Zhe Sun, Qiuguo Zhu