arXiv:2609.22684v1 Announce Type: cross
Abstract: Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA...
By Wenzhuo Li, Qiongfeng Shi, Yi Zhou
arXiv:2609.34792v2 Announce Type: replace
Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...
By Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi
MemBodied introduces a fixed‑size episodic memory for Vision‑Language‑Action models, comprising an associative state that tracks interactions across policy calls and an episode anchor that stores a compact representation of the initial scene. By conditioning action generation on these memory components instead of raw past observations, MemBodied reduces context bloat and inference latency. In five memory‑dependent RMBench tasks, it outperforms stateless and vanilla recurrent policies by significant margins, and achieves a 90.6% success rate on the LIBERO‑Long suite, improving over the baseline by 5.4%.
By Tej Deep Pala, Navonil Majumder, Bryce Goh, Raphael Yee, Jianfei Yang, Liming Chen, Soujanya Poria
arXiv:2603. 24576v2 Announce Type: replace-cross Abstract: Robots often observe information that determines a future action long before that action is executed.
By Xinying Guo, Chenxi Jiang, Hyun Bin Kim, Yuhang Han, Ying Sun, Yang Xiao, Jianfei Yang
arXiv:2610.00982v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the acti...
By Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh
arXiv:2610.00604v1 Announce Type: cross
Abstract: Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappe...
By Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov, Alexey K. Kovalev
arXiv:2609.36595v1 Announce Type: cross
Abstract: Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual...
By Yuyou Zhang, Yunbei Zhang, Miao Li, Janet Wang, Zijian Jin, Shilong Liu, Ding Zhao
arXiv:2607. 08032v1 Announce Type: new Abstract: Large language models, and the agents built on them, spend an ever-growing share of their compute and memory on remembering: caching attention keys and values, carrying long prompts, maintaining recurrent state, and storing what happened in previous turns and sessions.
By Ashwin Gerard Colaco, Nada Lahjouji
arXiv:2609.22854v1 Announce Type: cross
Abstract: Vision-language-action models often predict actions from only the current observation, which can leave tasks involving object occlusion or visually i...
By Jan-Gerrit Habekost, Parsa Mastouri Kashani, Connor G\"ade, Matthias Kerzel, Philipp Allgeuer, Cornelius Weber, Stefan Wermter, Jae Hee Lee
arXiv:2606. 20537v1 Announce Type: new Abstract: Mainstream LLM serving systems reuse prefix work mainly through paged or radix key-value (KV) caches.
By Liang Su
arXiv:2609.07047v1 Announce Type: cross
Abstract: Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Action policies are usually evaluated when t...
By Haiyang Sun, Haoxiao Wang, Junming Chen, Weicheng Fang, Zihao Su, Jingkun Yi, Wenyou Yi, Hao Chen, Zhou Zhao
arXiv:2609.32453v2 Announce Type: replace-cross
Abstract: Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a sho...
By Xinyu Zhao, Yixiang Shan, Tao Yang, Runyu Lei, Yiming Zhao, Jiaxin Fan, Zongbao Feng, Peng Jia