arXiv AI

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

arXiv:2606. 15768v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene.

arXiv Machine Learning
Jun 16

Geometric Action Model for Robot Policy Learning

arXiv:2606. 17046v1 Announce Type: cross Abstract: Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world.

By Jisang Han, Seonghu Jeon, Jaewoo Jung, Ren\'e Zurbr\"ugg, Honggyu An, Tifanny Portela, Marco Hutter, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
arXiv AI
Aug 27

DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation

DELE-w0.5 is a robotic manipulation framework that predicts future latent states instead of generating full video sequences, thereby inferring robot actions directly from these compact representations. By focusing on physical state changes rather than visual transitions, it reduces model complexity and inference latency. In 480 real‑robot trials across four long‑horizon tasks, DELE‑w0.5 achieved 62.5 % overall task success and 81.3 % macro ordered‑stage progress, outperforming the strongest baseline by 47.5 and 30.7 percentage points.

By Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren
arXiv Computer Vision
Sep 16

World-Action Models for Robot Learning and Control: A Survey

The survey "World-Action Models for Robot Learning and Control" reviews recent advances in coupling future world prediction with executable action generation for robots in open environments. It clarifies the scope of World-Action Models (WAMs) relative to conventional world models, model-based RL, and Vision‑Language‑Action policies, and organizes existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. The paper also surveys applications in manipulation, navigation, and autonomous driving, summarizes datasets, benchmarks, and metrics, and discusses key challenges such as action alignment, spatial consistency, long‑horizon memory, and efficient inference.

By Zuxing Lu, Hongjia Zhai, Guanzhi Wang, Huajian Zeng, Jiaqi Yang, Jingyu Liu, Lei Cheng, Yuantai Zhang, Yuheng Qiu, Zezhou Cheng, Ivan Laptev, Danfei Xu, Benjamin Riviere, Giuseppe Loianno, Eric Xing, Xingxing Zuo
arXiv Computer Vision
Sep 24

AWM-VLA: AlignedWorld Modeling for Efficient and Explainable Vision-Language-Action Policies

AWM‑VLA introduces a unified framework that embeds aligned world modeling directly into a diffusion‑transformer vision‑language‑action policy. By adding learnable future tokens aligned with vision‑language embeddings of future observations, the policy can anticipate long‑term consequences while generating actions. The method extends this with an object‑centric alignment objective and a principled weighting scheme, achieving up to 21% higher success rates on RoboCasa and humanoid tabletop benchmarks and producing object‑centric rationales preferred by human raters in 83% of cases.

By An Lanji, Dawei Liu, Jin Li, Haoran Xu, Mei Chen, Yu Tian
arXiv AI
Aug 28

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

The paper introduces a Latent World Model (LWM) for robot navigation that predicts action‑conditioned latent feature compatibility instead of reconstructing future observations. By exploiting the correlation between spatial proximity and latent feature similarity, the model evaluates action consequences directly in latent space and supports counterfactual training using sampled action sequences. The learned world model can supervise policy learning from unlabeled video and further improve policies via reinforcement learning entirely within the model, eliminating the need for action annotations and additional environment interaction.

By Zengmao Wang, Wei Gao, Shuhan Shen
arXiv AI
6d ago

Towards VLA-Dreamer: Refining VLA Behavior Using World Models

The paper proposes a new architecture for Vision‑Language‑Action (VLA) models that improves sample efficiency by training a predictive world model on the vision encoder’s embedding space. It argues that these embeddings are action‑relevant and can be used to predict future states, addressing the lack of an explicit world model in current VLAs. The trained model can also support short‑term planning by sampling actions that lead to desired goal images.

By Parsa Mastouri Kashani, Jan-Gerrit Habekost, Stefan Wermter