Latent evolving World Action Model
arXiv:2609.27455v1 Announce Type: new Abstract: World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). I...
arXiv:2609.27455v1 Announce Type: new Abstract: World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). I...
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
MT‑WAM enhances the Fast‑WAM framework by adding complementary supervision for future 2‑D point trajectories and visual features while keeping the original training objectives. A lightweight dual‑stream branch and structured attention mask isolate motion‑specific processing, and motion‑stream tokens provide additional dynamics cues to the action expert. During inference, MT‑WAM skips future‑video prediction, using cached video and motion information to achieve higher success rates on LIBERO, LIBERO‑Plus, RoboTwin 2.0 Clean2Rand, and several real‑world tasks.
JEPA-WAM enhances World Action Models (WAMs) by pairing text instructions with stochastically generated visual cues, using a text-to-image generator and a frozen V‑JEPA encoder to create dense goal representations. These representations are compressed into goal tokens that condition both video and action experts via cross‑attention, enabling the model to better ground instructions. On a new real‑robot benchmark, JEPA‑WAM attains 87.3%, 74.5%, and 80.9% success rates across in‑distribution, out‑of‑distribution scenes, and out‑of‑distribution instructions, outperforming prior methods by significant margins.
arXiv:2606. 15768v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene.
InternW0-Δ is a unified World Action Model that integrates pretrained visual dynamics, scene semantics, 4D geometry, and motion priors within a Mixture-of-Transformers framework to generate robot actions. It leverages a frozen VLM for semantic guidance, a 4D foundation model for geometric priors, and introduces Causal Imprint to learn future-relevant scene changes without future-video rollout. The model is pretrained on a newly curated 20K‑hour heterogeneous corpus of robot and human demonstrations, achieving superior performance on simulation benchmarks and real‑robot platforms.
arXiv:2609.23369v1 Announce Type: new Abstract: Generation-free world action models (WAMs) retain future-video prediction during training but act from internal video features at inference, leaving un...
arXiv:2607. 27138v1 Announce Type: cross Abstract: Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change.
arXiv:2606.27504v2 Announce Type: replace Abstract: World Action Models (WAMs) unify future environment prediction with action generation for autonomous driving, yet existing approaches optimize only...
arXiv:2607. 08182v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to map multimodal inputs to robot actions.
arXiv:2609.37165v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific f...
arXiv:2608.20974v1 Announce Type: cross Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent featu...