LT-Mem introduces a volatility‑aware memory evolution framework for lifelong scene understanding, combining spatially aligned instance‑level 3D perception with temporal reasoning. It uses a multi‑session SLAM backbone, a reasoning layer that scores evidence and selects memory actions, and a Tri‑Memory structure (Live, Delta, Meta) to preserve current states and event histories. The accompanying LT‑VQA dataset provides multi‑session recordings, persistent identity annotations, and temporal QA pairs, and experiments show LT‑Mem outperforms baselines while using far fewer tokens.
Mem-World introduces a memory‑augmented action‑conditioned world model for robot manipulation, featuring W‑VMem—a 4D wrist‑view‑centered surfel‑indexed memory that anchors historical observations to evolving surface elements. By explicitly modeling when and where scene elements are observed, the system retrieves geometry‑aware history frames during generation, providing informative, non‑redundant context for future action predictions. Experiments demonstrate that Mem‑World produces persistent rollouts, improves policy evaluation reliability (14.5 % higher Pearson correlation with real‑world performance), and boosts long‑horizon task success rates from 58 % to 72 % using synthetic data generation.
By Zirui Zheng, Jiaqian Yu, Xiongfeng Peng, jun shi, Mingyi Li, Chao Zhang, Weiming Li, Dong Wang, Huchuan Lu, Xu Jia
arXiv:2606. 20209v1 Announce Type: cross Abstract: Joint spatial and temporal understanding of 3D scenes is a crucial requirement for robots deployed in everyday household environments.
By Francesco Argenziano, Miguel Saavedra-Ruiz, Sacha Morin, Charlie Gauthier, Daniele Nardi, Liam Paull
The paper introduces UniWM, a unified, memory‑augmented world model that merges egocentric visual foresight and planning into a single multimodal autoregressive backbone. By grounding action selection in visually imagined outcomes and using a hierarchical memory to fuse short‑term perception with long‑term trajectory context, UniWM aligns prediction with control and improves navigation stability. Experiments on four challenging benchmarks and the 1X Humanoid Dataset show up to 30% higher success rates, reduced trajectory errors, zero‑shot generalization to unseen datasets, and scalability to high‑dimensional humanoid navigation.
By Yifei Dong, Fengyi Wu, Guangyu Chen, Lingdong Kong, Qiyu Hu, Yuxuan Zhou, Xu Zhu, Jingdong Sun, Jun-Yan He, Qi Dai, Alexander G. Hauptmann, Zhi-Qi Cheng
arXiv:2512. 21201v3 Announce Type: replace-cross Abstract: Zero-shot object navigation (ZSON) requires robots to find target objects in unseen environments without task-specific fine-tuning or pre-built maps, a key capability for general-purpose service robots.
By Yu He, Da Huang, Zhenyang Liu, Zixiao Gu, Qiang Sun, Guangnan Ye, Yanwei Fu, Yu-Gang Jiang
arXiv:2606. 18888v1 Announce Type: new Abstract: Navigation in partially observable environments presents a significant challenge for autonomous agents, requiring effective decision-making with limited sensory information in unknown environments.
By Thomas Quilter, Yifan Zhu, Guorui Quan, Mingfei Sun, Samuel Kaski