arXiv:2606. 18634v1 Announce Type: cross Abstract: To locate a target object while exploring the unknown environment is a fundamental capability for autonomous agents, with applications ranging from search-and-rescue to field robots.
By Zecheng Yin, Benedict Jun Ma
The paper introduces EvolvingNav, a system that builds a time‑indexed belief about moving targets in dynamic environments by combining timestamped 3D object histories with a persistence‑relocation model. It uses an event‑driven filter to update beliefs over time, incorporates RGB‑D evidence, and applies a zero‑shot vision‑language controller for action selection. The authors also present EvoWorld‑Bench, a large benchmark of human‑activity‑based scenes, and demonstrate that EvolvingNav outperforms baselines in both simulation and real‑robot experiments, especially when temporal patterns are learnable.
By Mingjian Gao, Zhaocheng Li, Haoyang Huang, Wenqiao Zhang, Yingjie Niu, Hao Zhou, Chao Li, Juncheng Li, Siliang Tang, Yueting Zhuang
arXiv:2608. 07079v1 Announce Type: cross Abstract: Object-goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi-object navigation and cross-floor navigation are still commonly addressed separately.
By Zehui Li, Zihao Sun, Jiawei Xu, Zheqi He, Xiaoqiang Zhang, Jing-Shu Zheng, Lu Liu, Dahui Gao, Xiuwan Chen
arXiv:2609.22351v1 Announce Type: new
Abstract: Open-vocabulary 3D Scene Graphs (3DSGs) ground each object node in a vision-language embedding, yet they record every entry as equally certain, so a ro...
By Carlos Cueto Zumaya, Iacopo Catalano, Wallace Moreira Bessa, Julio A. Placed
arXiv:2609.18058v1 Announce Type: new
Abstract: Finding the object referred to by language in a partially observed 3D scene is a core capability for embodied agents. Existing approaches either couple...
By Shixiong Xu, Zhiyuan Chen, Song Ding, Rui Luo, Xiaowei Liang, Dongxu Miao, Zhiying Du
arXiv:2603. 23800v2 Announce Type: replace-cross Abstract: We present a novel LLM-informed model-based planning framework, and a novel prompt selection method, for object search in partially-known environments.
By Abhishek Paudel, Abhish Khanal, Raihan I. Arnob, Shahriar Hossain, Gregory J. Stein
arXiv:2609.27076v1 Announce Type: new
Abstract: Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perce...
By Linus Nwankwo, Muslim Alaran, Christian Rauch, Stanley Chukwuebuka Obilikpa, Elmar Rueckert
arXiv:2606. 22338v2 Announce Type: replace-cross Abstract: Robots deployed in realistic settings will accumulate experience across many sessions and tasks over their deployment.
By Soumil Rathi
arXiv:2607. 14514v1 Announce Type: cross Abstract: Object-goal navigation requires an embodied agent to locate and reach an instance of a specified object category in an indoor environment.
By Xiaoran Xu, Yupeng Wu, Tianyu Xue, Yifan Xu, Xuanran Dong, Xiaoshan Yang, Changsheng Xu
LT-Mem introduces a volatility‑aware memory evolution framework for lifelong scene understanding, combining spatially aligned instance‑level 3D perception with temporal reasoning. It uses a multi‑session SLAM backbone, a reasoning layer that scores evidence and selects memory actions, and a Tri‑Memory structure (Live, Delta, Meta) to preserve current states and event histories. The accompanying LT‑VQA dataset provides multi‑session recordings, persistent identity annotations, and temporal QA pairs, and experiments show LT‑Mem outperforms baselines while using far fewer tokens.
arXiv:2609.23534v1 Announce Type: new
Abstract: Language-based 3D localization retrieves the point-cloud submap containing a target position from descriptions of nearby objects and their spatial rela...
By Tianyi Shang, Yike Shi, Zhenyu Li
arXiv:2609.22684v1 Announce Type: cross
Abstract: Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA...
By Wenzhuo Li, Qiongfeng Shi, Yi Zhou