arXiv:2610.07540v1 Announce Type: new
Abstract: Robotic systems often exhibit unstable modes, along which small perturbations and disturbances can cause unbounded growth unless corrected through feed...
By Leonardo F. Toso, Yann LeCun, James Anderson, Oumayma Bounou
arXiv:2511. 14427v4 Announce Type: replace-cross Abstract: Effective contact-rich manipulation requires robots to synergistically leverage vision, force, and proprioception.
By Rickmer Krohn, Vignesh Prasad, Gabriele Tiboni, Georgia Chalvatzaki
The paper presents a vision‑language navigation system that transfers from simulation to a real Ackermann‑steered mobile robot without relying on navigation graphs or panoramic views. It uses a Cross‑Modal Attention architecture trained on simulated data and fine‑tuned with limited real‑world episodes, leveraging linear photometric adjustments and a camera‑LiDAR sensor suite. Evaluation with SPL and nDTW metrics shows robust, adaptable navigation in continuous environments.
By Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara
arXiv:2608. 19866v1 Announce Type: new Abstract: This paper presents a novel data-driven approach to camera-based autonomy for micro-drones in GPS-denied, radio-challenging indoor environments.
By Niklas Voigt, Hartmut Surmann
PACT‑WAM is a world‑action model that simultaneously predicts a 16‑step action trajectory and its corresponding visual forecast for robot manipulation. It uses a hierarchical history encoder that compresses past observations into fewer tokens, reducing processing cost by 75% compared to dense encoding. The model’s shared flow module updates action and visual states jointly, and a TiTok‑VAE decoder reconstructs multi‑view future images, which are then used by a vision‑language component (Proposal Review) to improve execution‑prefix selection and proposal rejection, boosting success rates on several benchmarks.
By Yushan Liu, Jingjing Fan, Shoujie Li, Yifan Xie, Xiao-Ping Zhang, Wenbo Ding
Underwater C³-JEPA is an object‑centric, cross‑view predictive world model designed for near‑field heavy‑load ROV salvage. It encodes synchronized multi‑camera RGB observations and vehicle control signals into task‑object and context tokens, fuses cross‑camera evidence via held‑out‑view attention, and predicts future states conditioned on control without using contact sensors. The model demonstrates superior transfer of task‑relevant information to downstream probes compared to a reconstruction‑free baseline, supports model‑predictive control and imagined‑rollout training, and validates its effectiveness on real underwater video by accurately recovering withheld camera states and maintaining predictive lead over persistence.
By Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang
arXiv:2606. 16605v1 Announce Type: new Abstract: World models are widely used in robotic and agentic engineering control systems due to their ability to learn latent dynamics for planning and decision-making.
By Junjian Zhang, Hao Tan, Ruonan Li, Dong Zhu, Aiping Li, Zhaoquan Gu
The survey "World-Action Models for Robot Learning and Control" reviews recent advances in coupling future world prediction with executable action generation for robots in open environments. It clarifies the scope of World-Action Models (WAMs) relative to conventional world models, model-based RL, and Vision‑Language‑Action policies, and organizes existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. The paper also surveys applications in manipulation, navigation, and autonomous driving, summarizes datasets, benchmarks, and metrics, and discusses key challenges such as action alignment, spatial consistency, long‑horizon memory, and efficient inference.
By Zuxing Lu, Hongjia Zhai, Guanzhi Wang, Huajian Zeng, Jiaqi Yang, Jingyu Liu, Lei Cheng, Yuantai Zhang, Yuheng Qiu, Zezhou Cheng, Ivan Laptev, Danfei Xu, Benjamin Riviere, Giuseppe Loianno, Eric Xing, Xingxing Zuo
arXiv:2608. 05989v1 Announce Type: new Abstract: Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL).
By Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen
arXiv:2606. 05979v1 Announce Type: cross Abstract: We propose world-language-action (WLA) models as a new class of embodied foundation models.
By Yi Yang, Zhihong Liu, Siqi Kou, Yiyang Chen, Yanzhe Hu, Jianbo Zhou, Boyuan Zhao, Zhijie Wei, Xiao Xia, Xueqi Li, Pengfei Liu, Zhijie Deng
arXiv:2609.38163v1 Announce Type: new
Abstract: World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and pred...
By Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang
arXiv:2608. 20114v1 Announce Type: new Abstract: Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control.
By Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Qiaojun Yu