arXiv:2609.24124v1 Announce Type: cross
Abstract: Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effect...
By Yibo Li, Enshen Zhou, Rui Chen, Yanjun Ding, Mengzhen Liu, Yi Han, Jiabo Zhan, Lipeng Wang, Shanghang Zhang, Lu Sheng
The paper introduces 2AM, a system that separates memory and action execution in long‑horizon robot manipulation. 2AM stores task memory exclusively in a multimodal Agent, while a single RGB‑based, stateless Action Model performs motion based on language and optional 2D hints. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion—over 61 points higher than the best baseline—demonstrating that agent‑side memory and precise steering of the Action Model can substantially improve performance.
By Yutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry
The paper introduces 2AM, a system that keeps task memory solely within a multimodal Agent while using a single RGB‑based, stateless Action Model to execute motions. By compiling interaction history into subtask language and optional 2D grasp/place/move hints, the Agent steers the Action Model, which is trained to tolerate imperfect guidance through dropout, noise, and jitter. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion without depth, geometry, or planners, vastly outperforming the best baseline.
arXiv:2609.36595v1 Announce Type: cross
Abstract: Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual...
By Yuyou Zhang, Yunbei Zhang, Miao Li, Janet Wang, Zijian Jin, Shilong Liu, Ding Zhao
arXiv:2609.22684v1 Announce Type: cross
Abstract: Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA...
By Wenzhuo Li, Qiongfeng Shi, Yi Zhou
Mem-World introduces a memory‑augmented action‑conditioned world model for robot manipulation, featuring W‑VMem—a 4D wrist‑view‑centered surfel‑indexed memory that anchors historical observations to evolving surface elements. By explicitly modeling when and where scene elements are observed, the system retrieves geometry‑aware history frames during generation, providing informative, non‑redundant context for future action predictions. Experiments demonstrate that Mem‑World produces persistent rollouts, improves policy evaluation reliability (14.5 % higher Pearson correlation with real‑world performance), and boosts long‑horizon task success rates from 58 % to 72 % using synthetic data generation.
By Zirui Zheng, Jiaqian Yu, Xiongfeng Peng, jun shi, Mingyi Li, Chao Zhang, Weiming Li, Dong Wang, Huchuan Lu, Xu Jia
arXiv:2604. 18933v2 Announce Type: replace-cross Abstract: Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks that demand in-context memorization of historical information within a single trial or in-context adaptation based on the outcomes of multiple past trials.
By Yihuai Gao, Jeff Jinyun Liu, Shuang Li, Shuran Song
arXiv:2609.34792v2 Announce Type: replace
Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...
By Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi
arXiv:2603. 24576v2 Announce Type: replace-cross Abstract: Robots often observe information that determines a future action long before that action is executed.
By Xinying Guo, Chenxi Jiang, Hyun Bin Kim, Yuhang Han, Ying Sun, Yang Xiao, Jianfei Yang
The paper introduces Divide-and-Remember (D&R), a recursive memory method for vision-language-action (VLA) policies that optimises memory by maximizing the conditional mutual information between actions and memory given observations. D&R recursively divides the full history into top‑K selections over 2K tokens, using a shared lightweight selector across all recursion blocks to handle unbounded histories efficiently. Evaluated on the RoboMME benchmark of 16 long‑horizon manipulation tasks, D&R achieves state‑of‑the‑art success rates with consistent gains across all suites while using only 64 tokens, and similar improvements are observed in real‑robot experiments.
By Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh
LT-Mem introduces a volatility‑aware memory evolution framework for lifelong scene understanding, combining spatially aligned instance‑level 3D perception with temporal reasoning. It uses a multi‑session SLAM backbone, a reasoning layer that scores evidence and selects memory actions, and a Tri‑Memory structure (Live, Delta, Meta) to preserve current states and event histories. The accompanying LT‑VQA dataset provides multi‑session recordings, persistent identity annotations, and temporal QA pairs, and experiments show LT‑Mem outperforms baselines while using far fewer tokens.
arXiv:2606. 14551v1 Announce Type: cross Abstract: Robots under autonomous operation may require decisions based on evidence that is no longer visible.
By Zihao Li, Ranpeng Qiu, Yincong Chen, Guoqiang Ren, Weiming Zhi