arXiv Machine Learning
Aug 31

OceanGym: A Benchmark Environment for Underwater Embodied Agents

OceanGym is the first comprehensive benchmark for underwater embodied agents, featuring eight realistic task domains and a unified agent framework powered by Multi‑modal Large Language Models (MLLMs). It challenges agents to process optical and sonar data, navigate complex environments, and achieve long‑horizon goals amid low visibility and dynamic currents. Experiments show significant performance gaps between current MLLM agents and human experts, underscoring the difficulty of perception, planning, and adaptability in ocean settings.

By Yida Xue, Mingjun Mao, Xiangyuan Ru, Yuqi Zhu, Baochang Ren, Shuofei Qiao, Mengru Wang, Shumin Deng, Xinyu An, Ningyu Zhang, Ying Chen, Huajun Chen
arXiv AI
Sep 25

Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage

Underwater C³-JEPA is an object‑centric, cross‑view predictive world model designed for near‑field heavy‑load ROV salvage. It encodes synchronized multi‑camera RGB observations and vehicle control signals into task‑object and context tokens, fuses cross‑camera evidence via held‑out‑view attention, and predicts future states conditioned on control without using contact sensors. The model demonstrates superior transfer of task‑relevant information to downstream probes compared to a reconstruction‑free baseline, supports model‑predictive control and imagined‑rollout training, and validates its effectiveness on real underwater video by accurately recovering withheld camera states and maintaining predictive lead over persistence.

By Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang
arXiv Computer Vision
Sep 16

World-Action Models for Robot Learning and Control: A Survey

The survey "World-Action Models for Robot Learning and Control" reviews recent advances in coupling future world prediction with executable action generation for robots in open environments. It clarifies the scope of World-Action Models (WAMs) relative to conventional world models, model-based RL, and Vision‑Language‑Action policies, and organizes existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. The paper also surveys applications in manipulation, navigation, and autonomous driving, summarizes datasets, benchmarks, and metrics, and discusses key challenges such as action alignment, spatial consistency, long‑horizon memory, and efficient inference.

By Zuxing Lu, Hongjia Zhai, Guanzhi Wang, Huajian Zeng, Jiaqi Yang, Jingyu Liu, Lei Cheng, Yuantai Zhang, Yuheng Qiu, Zezhou Cheng, Ivan Laptev, Danfei Xu, Benjamin Riviere, Giuseppe Loianno, Eric Xing, Xingxing Zuo