arXiv AI

GaussMemory: Task-Driven 3D Gaussian Scene Memory for Long-Horizon Robotic Manipulation

arXiv:2608. 14986v1 Announce Type: cross Abstract: Long-horizon robotic manipulation fundamentally relies on persistent spatial memory.

arXiv AI
Sep 12

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

The paper introduces 2AM, a system that separates memory and action execution in long‑horizon robot manipulation. 2AM stores task memory exclusively in a multimodal Agent, while a single RGB‑based, stateless Action Model performs motion based on language and optional 2D hints. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion—over 61 points higher than the best baseline—demonstrating that agent‑side memory and precise steering of the Action Model can substantially improve performance.

By Yutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry
Hugging Face Trending Papers
Sep 10

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

The paper introduces 2AM, a system that keeps task memory solely within a multimodal Agent while using a single RGB‑based, stateless Action Model to execute motions. By compiling interaction history into subtask language and optional 2D grasp/place/move hints, the Agent steers the Action Model, which is trained to tolerate imperfect guidance through dropout, noise, and jitter. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion without depth, geometry, or planners, vastly outperforming the best baseline.

arXiv Computer Vision
Sep 16

Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation

Mem-World introduces a memory‑augmented action‑conditioned world model for robot manipulation, featuring W‑VMem—a 4D wrist‑view‑centered surfel‑indexed memory that anchors historical observations to evolving surface elements. By explicitly modeling when and where scene elements are observed, the system retrieves geometry‑aware history frames during generation, providing informative, non‑redundant context for future action predictions. Experiments demonstrate that Mem‑World produces persistent rollouts, improves policy evaluation reliability (14.5 % higher Pearson correlation with real‑world performance), and boosts long‑horizon task success rates from 58 % to 72 % using synthetic data generation.

By Zirui Zheng, Jiaqian Yu, Xiongfeng Peng, jun shi, Mingyi Li, Chao Zhang, Weiming Li, Dong Wang, Huchuan Lu, Xu Jia
arXiv AI
Aug 5

Gated Memory Policy: In-Context Memorization and Adaptation

arXiv:2604. 18933v2 Announce Type: replace-cross Abstract: Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks that demand in-context memorization of historical information within a single trial or in-context adaptation based on the outcomes of multiple past trials.

By Yihuai Gao, Jeff Jinyun Liu, Shuang Li, Shuran Song
arXiv Computer Vision
5d ago

D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

arXiv:2609.34792v2 Announce Type: replace Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...

By Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi
arXiv AI
4d ago

Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

The paper introduces Divide-and-Remember (D&R), a recursive memory method for vision-language-action (VLA) policies that optimises memory by maximizing the conditional mutual information between actions and memory given observations. D&R recursively divides the full history into top‑K selections over 2K tokens, using a shared lightweight selector across all recursion blocks to handle unbounded histories efficiently. Evaluated on the RoboMME benchmark of 16 long‑horizon manipulation tasks, D&R achieves state‑of‑the‑art success rates with consistent gains across all suites while using only 64 tokens, and similar improvements are observed in real‑robot experiments.

By Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh
Hugging Face Trending Papers
Aug 19

LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding

LT-Mem introduces a volatility‑aware memory evolution framework for lifelong scene understanding, combining spatially aligned instance‑level 3D perception with temporal reasoning. It uses a multi‑session SLAM backbone, a reasoning layer that scores evidence and selects memory actions, and a Tri‑Memory structure (Live, Delta, Meta) to preserve current states and event histories. The accompanying LT‑VQA dataset provides multi‑session recordings, persistent identity annotations, and temporal QA pairs, and experiments show LT‑Mem outperforms baselines while using far fewer tokens.