arXiv:2609.24124v1 Announce Type: cross
Abstract: Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effect...
By Yibo Li, Enshen Zhou, Rui Chen, Yanjun Ding, Mengzhen Liu, Yi Han, Jiabo Zhan, Lipeng Wang, Shanghang Zhang, Lu Sheng
The paper introduces 2AM, a system that separates memory and action execution in long‑horizon robot manipulation. 2AM stores task memory exclusively in a multimodal Agent, while a single RGB‑based, stateless Action Model performs motion based on language and optional 2D hints. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion—over 61 points higher than the best baseline—demonstrating that agent‑side memory and precise steering of the Action Model can substantially improve performance.
By Yutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry
The paper introduces 2AM, a system that keeps task memory solely within a multimodal Agent while using a single RGB‑based, stateless Action Model to execute motions. By compiling interaction history into subtask language and optional 2D grasp/place/move hints, the Agent steers the Action Model, which is trained to tolerate imperfect guidance through dropout, noise, and jitter. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion without depth, geometry, or planners, vastly outperforming the best baseline.
arXiv:2609.36595v1 Announce Type: cross
Abstract: Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual...
By Yuyou Zhang, Yunbei Zhang, Miao Li, Janet Wang, Zijian Jin, Shilong Liu, Ding Zhao
arXiv:2609.22684v1 Announce Type: cross
Abstract: Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA...
By Wenzhuo Li, Qiongfeng Shi, Yi Zhou
Mem-World introduces a memory‑augmented action‑conditioned world model for robot manipulation, featuring W‑VMem—a 4D wrist‑view‑centered surfel‑indexed memory that anchors historical observations to evolving surface elements. By explicitly modeling when and where scene elements are observed, the system retrieves geometry‑aware history frames during generation, providing informative, non‑redundant context for future action predictions. Experiments demonstrate that Mem‑World produces persistent rollouts, improves policy evaluation reliability (14.5 % higher Pearson correlation with real‑world performance), and boosts long‑horizon task success rates from 58 % to 72 % using synthetic data generation.
By Zirui Zheng, Jiaqian Yu, Xiongfeng Peng, jun shi, Mingyi Li, Chao Zhang, Weiming Li, Dong Wang, Huchuan Lu, Xu Jia