Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by introducing SPEAR: A Simulator for Photorealistic Embodied AI Research.
arXiv:2607. 21522v1 Announce Type: cross Abstract: Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging.
By Hongxin Zhang, Chunru Lin, Junyan Li, Zhou Xian, Tsun-Hsuan Wang, Chuang Gan
arXiv:2608.24212v1 Announce Type: new
Abstract: The advancement of Embodied AI necessitates high-quality simulation assets that faithfully mirror the real world. However, transforming raw visual obse...
By Yumeng He, Yichen Song, Xiaotian Yang, Weijia Zhang, Zanwei Zhou, Junru Gong, Xiaokang Yang, Yunbo Wang
4DSynth is a controllable procedural system that transforms natural-language descriptions, blueprint masks, or single photographs into editable 4D environments featuring explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation states. The system unifies animation, camera planning, rendering, and task generation within a single geometry-grounded representation, enabling scalable creation of dynamic embodied simulation scenes. Using 4DSynth, the authors built 4DSynth-Nav, an interactive navigation benchmark that demonstrates the reproducibility and tunability of procedural failures across vision‑language models.
By Zehao Qi, Haochen Luo, Jia-Wang Bian, Zeyu Ma, Shuyang Sun
arXiv:2606. 18363v1 Announce Type: cross Abstract: Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents.
By Haowen Liu, Xirui Li, Shaoxiong Yao, Peng Shi, Tianyi Zhou, Jia-Bin Huang, Furong Huang, Jiayuan Mao
Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situations, and interacting with the physical world. In this report, we introduce Hy-Embodied-VLM-1.
The advancement of Embodied AI necessitates high-quality simulation assets that faithfully mirror the real world. However, transforming raw visual observations into simulation-ready scenes remains cha...
FluxVLA Engine is an open, configuration‑driven platform that unifies the fragmented components of embodied policy development—datasets, visual‑language and world models, action heads, learning methods, distributed training, simulation evaluation, inference, and robot interfaces—into a reproducible data‑to‑deployment workflow. It adds features such as compositional dual‑arm simulation, scalable automatic data generation, human‑in‑the‑loop rollout and correction, Real‑Time Chunking for fast inference, and lightweight remote GPU serving, thereby linking offline learning, simulation validation, online correction, and real‑robot execution under shared, auditable contracts. The engine aims to eliminate engineering bottlenecks that currently separate promising embodied‑learning algorithms from reliable, reproducible deployment.
By Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong, Qirui Hu, Yiming Zhang, Weipeng Deng, Bowen Shen, Minzhao Zhu, Yiming Mao, Yan Yang, Chenguang Cui, Hongyuan Chen, Xu Huang, Zheyi Zhao, Pinxi Shen, Bozhen He, Zhen Fu, Yifan Wang, Zexin Zhang, Ang Gao, Haoyu Chen, Chengqi Shi, Hua Chen
arXiv:2607. 11643v1 Announce Type: cross Abstract: Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints.
By Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
arXiv:2609.25627v1 Announce Type: cross
Abstract: General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate prec...
By Haoran Wen, Wenfu Wang, Kunsong Shi, Jingke Wang, Wancheng Feng, Yiren Zhang, Yueran Zhao, Xuancheng Zhang, Nanfei Ye, Xingru Chen, Zhaohong Sun, Chengmin Yang, Zikang Yu, Penghao Bi, Jia Shi, Yu Liu, Kun Zhan, Yan Xie
arXiv:2606.18363v3 Announce Type: replace-cross
Abstract: Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through...
By Haowen Liu, Xirui Li, Shaoxiong Yao, Peng Shi, Tianyi Zhou, Jia-Bin Huang, Furong Huang, Jiayuan Mao
arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.
By Yifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao, Yucheng Hu, Qiyu Wu, Yuxiao Li, Zibin Dong, Fei Ni, Yan Zheng, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye Hao