arXiv Machine Learning

Coupled Control and Wireless World Models for Resilient Remote Robotic Control

The paper introduces a resilient remote robotic control framework that couples control and wireless world models using a Joint Embedding Predictive Architecture (JEPA). By learning latent representations from visual observations and radio frequency (RF) data, the system predicts future robot states and wireless conditions to schedule uplink transmissions efficiently. An adaptive resilience mechanism adjusts perception embeddings when prediction errors arise, enabling robust operation without retraining the entire control policy.

arXiv AI
1d ago

Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot

The paper presents a vision‑language navigation system that transfers from simulation to a real Ackermann‑steered mobile robot without relying on navigation graphs or panoramic views. It uses a Cross‑Modal Attention architecture trained on simulated data and fine‑tuned with limited real‑world episodes, leveraging linear photometric adjustments and a camera‑LiDAR sensor suite. Evaluation with SPL and nDTW metrics shows robust, adaptable navigation in continuous environments.

By Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara
arXiv AI
Sep 17

PACT-WAM: Predicting Actions and Visual Foresight with Compact Temporal Encoding for Robot Manipulation

PACT‑WAM is a world‑action model that simultaneously predicts a 16‑step action trajectory and its corresponding visual forecast for robot manipulation. It uses a hierarchical history encoder that compresses past observations into fewer tokens, reducing processing cost by 75% compared to dense encoding. The model’s shared flow module updates action and visual states jointly, and a TiTok‑VAE decoder reconstructs multi‑view future images, which are then used by a vision‑language component (Proposal Review) to improve execution‑prefix selection and proposal rejection, boosting success rates on several benchmarks.

By Yushan Liu, Jingjing Fan, Shoujie Li, Yifan Xie, Xiao-Ping Zhang, Wenbo Ding
arXiv AI
Sep 25

Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage

Underwater C³-JEPA is an object‑centric, cross‑view predictive world model designed for near‑field heavy‑load ROV salvage. It encodes synchronized multi‑camera RGB observations and vehicle control signals into task‑object and context tokens, fuses cross‑camera evidence via held‑out‑view attention, and predicts future states conditioned on control without using contact sensors. The model demonstrates superior transfer of task‑relevant information to downstream probes compared to a reconstruction‑free baseline, supports model‑predictive control and imagined‑rollout training, and validates its effectiveness on real underwater video by accurately recovering withheld camera states and maintaining predictive lead over persistence.

By Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang
arXiv Computer Vision
Sep 16

World-Action Models for Robot Learning and Control: A Survey

The survey "World-Action Models for Robot Learning and Control" reviews recent advances in coupling future world prediction with executable action generation for robots in open environments. It clarifies the scope of World-Action Models (WAMs) relative to conventional world models, model-based RL, and Vision‑Language‑Action policies, and organizes existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. The paper also surveys applications in manipulation, navigation, and autonomous driving, summarizes datasets, benchmarks, and metrics, and discusses key challenges such as action alignment, spatial consistency, long‑horizon memory, and efficient inference.

By Zuxing Lu, Hongjia Zhai, Guanzhi Wang, Huajian Zeng, Jiaqi Yang, Jingyu Liu, Lei Cheng, Yuantai Zhang, Yuheng Qiu, Zezhou Cheng, Ivan Laptev, Danfei Xu, Benjamin Riviere, Giuseppe Loianno, Eric Xing, Xingxing Zuo
arXiv Computer Vision
Sep 30

Rethinking Representations for World-Action Modeling

arXiv:2609.38163v1 Announce Type: new Abstract: World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and pred...

By Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang