OceanGym is the first comprehensive benchmark for underwater embodied agents, featuring eight realistic task domains and a unified agent framework powered by Multi‑modal Large Language Models (MLLMs). It challenges agents to process optical and sonar data, navigate complex environments, and achieve long‑horizon goals amid low visibility and dynamic currents. Experiments show significant performance gaps between current MLLM agents and human experts, underscoring the difficulty of perception, planning, and adaptability in ocean settings.
By Yida Xue, Mingjun Mao, Xiangyuan Ru, Yuqi Zhu, Baochang Ren, Shuofei Qiao, Mengru Wang, Shumin Deng, Xinyu An, Ningyu Zhang, Ying Chen, Huajun Chen
arXiv:2606. 13817v1 Announce Type: cross Abstract: World models in robot learning predict future states from visual observations and actions, enabling agents to reason about the consequences of their controls.
By Yitao Jiang, Luyang Zhao, Muhao Chen, Devin Balkcom
Underwater C³-JEPA is an object‑centric, cross‑view predictive world model designed for near‑field heavy‑load ROV salvage. It encodes synchronized multi‑camera RGB observations and vehicle control signals into task‑object and context tokens, fuses cross‑camera evidence via held‑out‑view attention, and predicts future states conditioned on control without using contact sensors. The model demonstrates superior transfer of task‑relevant information to downstream probes compared to a reconstruction‑free baseline, supports model‑predictive control and imagined‑rollout training, and validates its effectiveness on real underwater video by accurately recovering withheld camera states and maintaining predictive lead over persistence.
By Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang
The survey "World-Action Models for Robot Learning and Control" reviews recent advances in coupling future world prediction with executable action generation for robots in open environments. It clarifies the scope of World-Action Models (WAMs) relative to conventional world models, model-based RL, and Vision‑Language‑Action policies, and organizes existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. The paper also surveys applications in manipulation, navigation, and autonomous driving, summarizes datasets, benchmarks, and metrics, and discusses key challenges such as action alignment, spatial consistency, long‑horizon memory, and efficient inference.
By Zuxing Lu, Hongjia Zhai, Guanzhi Wang, Huajian Zeng, Jiaqi Yang, Jingyu Liu, Lei Cheng, Yuantai Zhang, Yuheng Qiu, Zezhou Cheng, Ivan Laptev, Danfei Xu, Benjamin Riviere, Giuseppe Loianno, Eric Xing, Xingxing Zuo
arXiv:2608. 20114v1 Announce Type: new Abstract: Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control.
By Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Qiaojun Yu
arXiv:2606. 05979v1 Announce Type: cross Abstract: We propose world-language-action (WLA) models as a new class of embodied foundation models.
By Yi Yang, Zhihong Liu, Siqi Kou, Yiyang Chen, Yanzhe Hu, Jianbo Zhou, Boyuan Zhao, Zhijie Wei, Xiao Xia, Xueqi Li, Pengfei Liu, Zhijie Deng
arXiv:2606. 08513v1 Announce Type: cross Abstract: Autonomous Underwater Vehicles (AUVs) traditionally rely on complex, heavily engineered pipelines for perception, path planning, and motion control.
By Elisei Shafer, Oren Gal
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
By Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang
CtrlWAM introduces a controllable world action model that jointly predicts actions (intent) and visual futures (foresight). By executing perturbed actions in a simulator and pairing them with noised visual outcomes, it aligns action predictions with their visual consequences, using warped video–action noise schedules to maintain visual layout responsiveness. The model extends beyond ego‑only control to multiple agent streams, improving action forecasts, video–action agreement, and command following in driving and robotics experiments.
By Chensheng Peng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, Masayoshi Tomizuka, Boris Ivanovic, Marco Pavone, Yuxiao Chen
World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.
arXiv:2608. 09298v1 Announce Type: cross Abstract: Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation.
By Peterson Co, Sicheng Hu, Chunxuan Jiao, Hongyang Cheng, Yulin Luo, Yijie Xu, Sixiang Chen, Zhongxia Zhao, Zihao Wang, DaFeng Chi, Peidong Liu, YuTong Chen, Henghua Liu, Zhihao Yuan, Huizhu Jia, Yuzheng Zhuang, Tianle Zhang, Liang Lin, Huajie Tan, Shanghang Zhang
arXiv:2609.38057v1 Announce Type: new
Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action m...
By Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian