arXiv:2607. 04171v3 Announce Type: replace-cross Abstract: Tiny Vision-Language-Action models are appealing for real-time robotic control, but reducing model scale often weakens two capabilities essential for manipulation: task-conditioned spatial grounding and coherent action generation.
By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
arXiv:2607.04171v4 Announce Type: replace-cross
Abstract: How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training fra...
By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
arXiv:2609.08638v1 Announce Type: cross
Abstract: An action chunk can span several stages of a manipulation task, yet a label for its first step describes only the current stage. We introduce Chunk-A...
By Tinghe Ding, Jiahao Li, He Wang
arXiv:2608.22364v1 Announce Type: new
Abstract: World action models (WAMs) couple visual future prediction with robot action generation, but accelerated students can lose task capabilities during dis...
By Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang
arXiv:2606. 25800v1 Announce Type: new Abstract: Effective online adaptation of vision-language-action (VLA) models remains challenging, as sparse rewards provide weak supervision for high-dimensional autoregressive action policies.
By Kejing Wang, Toan Nguyen, Minh Hoang Nguyen, Simon Khan, Flora D. Salim
arXiv:2607. 04171v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong multimodal understanding and spatial grounding, but their computational cost limits real-time robotic control.
By Lei Iok Tong, Qingchen Xie, Wei Huang, Ying Jie Yap, Yujie Zhang, Qianzhi Li, Xiaolong Liu, Zhidong Deng