arXiv:2606. 03598v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved remarkable success in language-conditioned robotic manipulation.
By Ziyang Chen, Shaoguang Wang, Weiyu Guo, Qianyi Cai, He Zhang, Pengteng Li, Yiren Zhao, Yandong Guo
Rewind-IL is a training‑free online safeguard for generative action‑chunked imitation learning policies. It uses a zero‑shot failure detector based on Temporal Inter‑chunk Discrepancy Estimate (TIDE) and a state‑respawning mechanism that returns the robot to a verified safe intermediate state. The system builds a checkpoint library offline with a vision‑language model and monitors self‑consistency online, rewinding execution to the latest safe checkpoint when a failure is detected, thereby improving reliability in long‑horizon manipulation tasks.
By Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson, Weiming Zhi
The paper investigates how to balance retaining past experience versus learning from new data when robot dynamics change. It introduces two metrics—change magnitude and age‑staleness AUC—to quantify when older transitions are helpful or harmful. Experiments on locomotion tasks and real‑world perturbations show that the optimal replay strategy depends on the size of the dynamics shift and the evolution of the system over time.
By Everest Yang, Skye Thompson, George D. Konidaris
The paper introduces a compositional continual learning benchmark for world models in robot manipulation, designed to isolate knowledge reuse from learning speed and capacity. Tasks are curated to combine previously seen action and perception components, allowing analysis of how different modalities affect reuse. Experiments show that modular world models better balance reuse and forgetting than conventional methods, yet none fully solve the challenge, highlighting the need for models explicitly built to reuse knowledge without forgetting.
By Haoyu Zhou, Joe Watson, Anson Lei, Ingmar Posner
arXiv:2609.35575v2 Announce Type: replace-cross
Abstract: The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrat...
By Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu, Chen Bai, Ziheng Zhang, Yufei Jia, Jindou Jia, Yuhang Zhang, Xinrui Zhang, Shang Yujing, Yuxiang Chen, Chuhao Zhou, Tiancai Wang, Jianfei Yang
arXiv:2603. 11395v3 Announce Type: replace-cross Abstract: Continual reinforcement learning challenges agents to acquire new skills while retaining previously learned ones with the goal of improving performance in both past and future tasks.
By Abdulaziz Alyahya, Abdallah Al Siyabi, Markus R. Ernst, Luke Yang, Levin Kuhlmann, Gideon Kowadlo
arXiv:2609.38057v1 Announce Type: new
Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action m...
By Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian
arXiv:2607. 04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks.
By Mao-Lin Luo, Zhe-Xu Wang, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
arXiv:2608. 07065v1 Announce Type: cross Abstract: Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single-step commands.
By Jinhe Tang, Weiming Zhi
arXiv:2607. 01111v1 Announce Type: cross Abstract: Robot policies inevitably encounter failures when deployed in real environments.
By Haoran Hao, Shahram Najam Syed, Jeffrey Ichnowski, Jeff Schneider
arXiv:2603. 11653v2 Announce Type: replace Abstract: Continual Reinforcement Learning (CRL) for Vision-Language-Action (VLA) models is a promising direction toward self-improving embodied agents that can adapt in openended, evolving environments.
By Jiaheng Hu, Jay Shim, Chen Tang, Yoonchang Sung, Bo Liu, Peter Stone, Roberto Martin-Martin
arXiv:2608. 15680v1 Announce Type: cross Abstract: Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states.
By Yijie Xu, Haopeng Jin, Run Zhou, Shengbang Liu, Sixiang Chen, Hongyang Cheng, Sicheng Hu, Peterson Co, Jinwen Luo, Huajie Tan, Shanghang Zhang