arXiv Computer Vision

What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling

arXiv AI
Sep 10

WorldAgen: Unified State-Action Prediction with Test-Time World Model Training

WorldAgen is a unified framework that jointly learns world modeling and action prediction using a shared Transformer backbone with two specialized heads. It introduces a Mixed Unidirectional Attention Mask to separate the world model and agent model, and enables Test-Time Training (TTT) by sampling exploratory actions and updating the world model with real state transitions. Experiments on CALVIN and LIBERO show that WorldAgen matches or surpasses state‑of‑the‑art methods, especially when TTT is applied to a few samples.

By Chi Wan, Kangrui Wang, Yuan Si, Pingyue Zhang, Manling Li
arXiv Computer Vision
1d ago

CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight

CtrlWAM introduces a controllable world action model that jointly predicts actions (intent) and visual futures (foresight). By executing perturbed actions in a simulator and pairing them with noised visual outcomes, it aligns action predictions with their visual consequences, using warped video–action noise schedules to maintain visual layout responsiveness. The model extends beyond ego‑only control to multiple agent streams, improving action forecasts, video–action agreement, and command following in driving and robotics experiments.

By Chensheng Peng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, Masayoshi Tomizuka, Boris Ivanovic, Marco Pavone, Yuxiao Chen
arXiv Machine Learning
Jun 9

C$^3$ache: Accelerating World Action Models with Cross Inference Chunk Cache

arXiv:2606. 08962v1 Announce Type: new Abstract: World Action Models (WAMs) generalize better than standard Vision-Language-Action (VLA) policies to novel motions and environments, because a video-modeling objective lets them learn from abundant unlabeled video rather than scarce labeled robot demonstrations.

By Weisen Zhao, Lam Nguyen, Zhicong Lu, Yuzhang Shang
arXiv AI
Sep 25

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

AD-WM is a new action‑discriminative joint‑embedding world model designed for counterfactual model predictive control. It augments residual latent dynamics with action‑recovery regularization based on inverse dynamics and conditional mutual information, while discarding auxiliary heads at test time so that MPC remains unchanged. Experiments on OGBench‑Cube and other simulation environments show substantial gains in hard‑start success and mean success, and zero‑shot transfer to a Franka robot improves pick‑and‑place success from 42.2% to 71.1%.

By Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo, Yang Gao