MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.22854v1 Announce Type: cross Abstract: Vision-language-action models often predict actions from only the current observation, which can leave tasks involving object occlusion or visually i...
arXiv:2609.34792v2 Announce Type: replace Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...
arXiv:2610.00982v1 Announce Type: cross Abstract: Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the acti...
arXiv:2607. 25487v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets.
arXiv:2603. 24576v2 Announce Type: replace-cross Abstract: Robots often observe information that determines a future action long before that action is executed.
arXiv:2608.29537v1 Announce Type: cross Abstract: Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so t...