arXiv:2607. 00836v1 Announce Type: cross Abstract: World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities.
By Xiaoxiong Zhang, Xiong Zeng, Wei Zhang
arXiv:2606. 17046v1 Announce Type: cross Abstract: Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world.
By Jisang Han, Seonghu Jeon, Jaewoo Jung, Ren\'e Zurbr\"ugg, Honggyu An, Tifanny Portela, Marco Hutter, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
arXiv:2606.27504v2 Announce Type: replace
Abstract: World Action Models (WAMs) unify future environment prediction with action generation for autonomous driving, yet existing approaches optimize only...
By Tianze Xia, Lijun Zhou, Kaixin Xiong, Jingfeng Yao, Zhenxin Zhu, Haiyang Sun, Bing Wang, Guang Chen, Wenyu Liu, Hangjun Ye, Xinggang Wang
The survey "World-Action Models for Robot Learning and Control" reviews recent advances in coupling future world prediction with executable action generation for robots in open environments. It clarifies the scope of World-Action Models (WAMs) relative to conventional world models, model-based RL, and Vision‑Language‑Action policies, and organizes existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. The paper also surveys applications in manipulation, navigation, and autonomous driving, summarizes datasets, benchmarks, and metrics, and discusses key challenges such as action alignment, spatial consistency, long‑horizon memory, and efficient inference.
By Zuxing Lu, Hongjia Zhai, Guanzhi Wang, Huajian Zeng, Jiaqi Yang, Jingyu Liu, Lei Cheng, Yuantai Zhang, Yuheng Qiu, Zezhou Cheng, Ivan Laptev, Danfei Xu, Benjamin Riviere, Giuseppe Loianno, Eric Xing, Xingxing Zuo
arXiv:2606. 15768v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene.
By Jialei Chen, Kai Wang, Kang Chen, Shuaihang Chen, Feng Gao, Wenhao Tang, Zhiyuan Li, Weilin Liu, Zhuyu Yao, Boxun Li, Yuanbo Xu, Chao Yu
The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.
By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
The paper introduces a Latent World Model (LWM) for robot navigation that predicts action‑conditioned latent feature compatibility instead of reconstructing future observations. By exploiting the correlation between spatial proximity and latent feature similarity, the model evaluates action consequences directly in latent space and supports counterfactual training using sampled action sequences. The learned world model can supervise policy learning from unlabeled video and further improve policies via reinforcement learning entirely within the model, eliminating the need for action annotations and additional environment interaction.
By Zengmao Wang, Wei Gao, Shuhan Shen
World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or...
arXiv:2609.38059v1 Announce Type: cross
Abstract: Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a sca...
By Shenghe Zheng, Wenbo Li, Jiyao Zhang, Bin Xia, Haoyang Huang, Nan Duan, Jiaya Jia
arXiv:2609.38057v1 Announce Type: new
Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action m...
By Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian
The Representation World Model (RWM) learns states, transitions, and executable plans directly within a representation space, bypassing traditional explicit dynamics models and action-space search. It uses inverse-dynamics supervision along latent paths to shape the representation geometry, enabling direct planning by constructing a latent path between current and goal states and recovering actions via inverse dynamics. Experiments on continuous-control benchmarks and robotic manipulation tasks demonstrate RWM’s effectiveness and potential for complex embodied control.
By Yijun Yuan, Weicheng Zheng, Weibang Wang, Minghui Qin, Chang Sun, Junhao Huang, Kenan Li, Anmin Liu, Yicheng Yao, Hang Zhao
arXiv:2609.10506v1 Announce Type: cross
Abstract: Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However...
By Nisarga Nilavadi, Ralf R\"omer, Moritz Reuss, Michael Krawez, Tobias J\"ulg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard