Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 19589v1 Announce Type: cross Abstract: Pretrained Vision-Language-Action models provide a strong foundation for robot learning, but sequentially adapting them to diverse skills can perturb the representations and velocity mappings used by previous skills, leading to catastrophic forgetting.
arXiv:2606. 03598v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved remarkable success in language-conditioned robotic manipulation.
arXiv:2609.32453v2 Announce Type: replace-cross Abstract: Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a sho...
arXiv:2609.22684v1 Announce Type: cross Abstract: Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA...
arXiv:2609.37889v1 Announce Type: cross Abstract: Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks whil...
arXiv:2610.00982v1 Announce Type: cross Abstract: Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the acti...