Mem2Ego introduces a vision‑language model for embodied navigation that combines global memory with egocentric visual inputs. By adaptively retrieving task‑relevant cues from a global memory module and aligning them with local perception, the framework improves spatial reasoning and decision‑making over long horizons. The method outperforms prior state‑of‑the‑art approaches on the HSSD and HM3D benchmarks and shows strong performance on a real robot.
By Lingfeng Zhang, Yuecheng Liu, Zhanguang Zhang, Matin Aghaei, Yixin Xiao, Yaochen Hu, Mohammad Ali Alomrani, David Gamaliel Arcos Bravo, Hongjian Gu, Zhiyuan Li, Yangzheng Wu, Zhanpeng Zhang, Raika Karimi, Atia Hamidizadeh, Guowei Huang, Haoping Xu, Tongtong Cao, Weichao Qiu, Xingyue Quan, Jianye Hao, Yuzheng Zhuang, Yingxue Zhang
The paper introduces EvolvingNav, a system that builds a time‑indexed belief about moving targets in dynamic environments by combining timestamped 3D object histories with a persistence‑relocation model. It uses an event‑driven filter to update beliefs over time, incorporates RGB‑D evidence, and applies a zero‑shot vision‑language controller for action selection. The authors also present EvoWorld‑Bench, a large benchmark of human‑activity‑based scenes, and demonstrate that EvolvingNav outperforms baselines in both simulation and real‑robot experiments, especially when temporal patterns are learnable.
By Mingjian Gao, Zhaocheng Li, Haoyang Huang, Wenqiao Zhang, Yingjie Niu, Hao Zhou, Chao Li, Juncheng Li, Siliang Tang, Yueting Zhuang
arXiv:2510. 14357v2 Announce Type: replace-cross Abstract: Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily relying on manual operations or fixed railways for movement.
By Xiaobei Zhao, Xingqi Lyu, Xin Chen, Xiang Li
Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored.
arXiv:2608. 14986v1 Announce Type: cross Abstract: Long-horizon robotic manipulation fundamentally relies on persistent spatial memory.
By Zhiqiang Hu, Shouren Huang, Masatoshi Ishikawa
arXiv:2610.00330v1 Announce Type: cross
Abstract: Long-term object search requires learning where objects usually appear from repeated but uneven observations of a changing environment. We formulate...
By Jiaming Wang, Zhiwei Xue, Chen Jizhuo, Peng Shiqi, Harold Soh
Researchers combined an efficient algorithm with dedicated hardware to rapidly generate 3D maps for navigation using minimal memory and power.
By Adam Zewe | MIT News
arXiv:2609.36595v1 Announce Type: cross
Abstract: Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual...
By Yuyou Zhang, Yunbei Zhang, Miao Li, Janet Wang, Zijian Jin, Shilong Liu, Ding Zhao
arXiv:2606. 08992v1 Announce Type: cross Abstract: Vision-and-Language Navigation in continuous environments requires agents to understand the spatial structure of previously unseen environments in order to follow language instructions.
By Yucheng Deng, Pingrui Lai, Xinhai Li, Chenjia Bai, Xiaoheng Deng, Chengnuo Sun, Xuelong Li, Hua Yang
LT-Mem introduces a volatility‑aware memory evolution framework for lifelong scene understanding, combining spatially aligned instance‑level 3D perception with temporal reasoning. It uses a multi‑session SLAM backbone, a reasoning layer that scores evidence and selects memory actions, and a Tri‑Memory structure (Live, Delta, Meta) to preserve current states and event histories. The accompanying LT‑VQA dataset provides multi‑session recordings, persistent identity annotations, and temporal QA pairs, and experiments show LT‑Mem outperforms baselines while using far fewer tokens.
arXiv:2608.31022v1 Announce Type: new
Abstract: AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However,...
By Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria
Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is difficult for real-world applications.