arXiv:2607. 18042v1 Announce Type: cross Abstract: End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly train the policy state to be predictive of future visual outcomes.
By Lingfeng Zhang, Zhanguang Zhang, Liheng Ma, Tongtong Cao, Yingxue Zhang
arXiv:2609.37250v1 Announce Type: cross
Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
By Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
By Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang
arXiv:2609.36645v1 Announce Type: cross
Abstract: Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encou...
By Hanseul Kim, Jewon Yeom, Youngjoon Jeong, Minsoo Jo, Taesup Kim
arXiv:2609.39324v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. Howeve...
By Jingqiu Wang, Yan Wang
arXiv:2606. 15768v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene.
By Jialei Chen, Kai Wang, Kang Chen, Shuaihang Chen, Feng Gao, Wenhao Tang, Zhiyuan Li, Weilin Liu, Zhuyu Yao, Boxun Li, Yuanbo Xu, Chao Yu