arXiv AI

MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation

arXiv Computer Vision
3d ago

D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

arXiv:2609.34792v2 Announce Type: replace Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...

By Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi
arXiv AI
Jul 29

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

arXiv:2607. 25487v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets.

By Minhyeok Lee, Chiyoung Kim, Chanhoe Gu, Seongrok Kim, Sanghyuk Roy Choi, Donghwan Hwang, Donghun Ryu, Seokhyun Kim
arXiv AI
Sep 24

MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

MemBodied introduces a fixed‑size episodic memory for Vision‑Language‑Action models, comprising an associative state that tracks interactions across policy calls and an episode anchor that stores a compact representation of the initial scene. By conditioning action generation on these memory components instead of raw past observations, MemBodied reduces context bloat and inference latency. In five memory‑dependent RMBench tasks, it outperforms stateless and vanilla recurrent policies by significant margins, and achieves a 90.6% success rate on the LIBERO‑Long suite, improving over the baseline by 5.4%.

By Tej Deep Pala, Navonil Majumder, Bryce Goh, Raphael Yee, Jianfei Yang, Liming Chen, Soujanya Poria
arXiv Machine Learning
Aug 19

VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation

VLCP (Vision Language Control Policy) is a training‑free robot manipulation approach that keeps a vision‑language model (VLM) frozen and uses it to generate short Python control functions. Unlike traditional methods that retry a fixed policy, VLCP rewrites the control code every K steps based on multi‑view RGB, proprioceptive state, and state delta, allowing failures to be corrected within the same episode. In a 57‑task MuJoCo/RoboVerse benchmark, VLCP achieves 35.1% pooled success versus 3.5% for a single‑query baseline, with a 27.3% within‑episode recovery rate on failed grasps and efficient token usage.

By Dhia Naouali, Minghan Wu, Claudia Wong, Abhinav Puthran, Omar G. Younis
arXiv AI
Aug 26

PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

PonderPounce introduces a two‑stage system that leverages a multimodal large language model (MLLM) as an episode‑context engine for robot control. The System2 component, Ponder, accumulates observations, demonstrations, and prior cognition in its native causal context, producing subgoal text and reasoning. The System1 component, Pounce, uses the current observation, instruction, and proprioception, receiving only the newest cognition token and its age from Ponder; the pair is jointly trained end‑to‑end without a separate memory module, achieving real‑time action playback and outperforming baseline methods on RoboMME and RoboCasa‑DC benchmarks.

By Suhwan Choi, Jaeyoon Jung, Sungkyung Kim, Yunsung Lee, Youngjae Yu