arXiv Machine Learning

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

arXiv:2607. 28362v1 Announce Type: cross Abstract: We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models.

arXiv AI
Aug 28

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

CLAP is a cross-embodiment framework for action‑conditioned video generation that can be trained on diverse internet‑scale videos from both humans and robots. It reconciles different action spaces—end‑effector poses, language instructions, and latent actions—using a curriculum that first learns physics priors from unlabeled video and then grounds them in real‑world action spaces for zero‑shot deployment. The resulting models match or exceed state‑of‑the‑art single‑embodiment models in challenging environments and support few‑shot adaptation across a wide range of robot morphologies.

By Kechen Liu, Ola Shorinwa
arXiv AI
Aug 28

GameWAM: A World Action Model for Video Games

GameWAM is the first World-Action Model designed for native closed-loop gameplay and GUI control in modern video games. It jointly generates future visual observations and executable keyboard-mouse trajectories using parallel visual and action generative processes, block-causal conditioning, and flow matching. The model predicts gameplay/GUI mode at each step, handles heterogeneous native controls, and employs block-cycle control for long-horizon interaction, achieving competitive task success with fewer native actions than prior agents.

By Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li
arXiv AI
6d ago

DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

DyMD introduces a Distribution Matching Distillation framework that adapts teacher supervision and critic fitting to preserve interaction dynamics in few-step video generation. By employing temporal affinity–conditioned re‑noise sampling and dynamics‑guided fake‑score tracking, DyMD balances motion recovery with visual quality. The method distills a 14B teacher into a 1.3B student that achieves significant gains on embodied‑video benchmarks and downstream action planning tasks.

By Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu
arXiv AI
6d ago

Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases

The paper introduces Action Forcing, a method that transforms ordinary unlabeled video into action‑supervised training data by extracting egomotion bases through principal component analysis of pixel displacements. This approach yields grounded throttle–yaw control signals without requiring instrumented platforms or manual annotation, and it trains a high‑capacity video model while preventing pixel‑level overfitting via an online latent critic. The authors also critique standard video generation metrics and propose a reference‑free evaluation that measures controllability, plausibility, conjuring, and geometric integrity, showing that their model can reverse, scale, and compose actions despite limited reverse‑action data.

By Ashish Sundar, Tiankuo Hou, Zhong Fan, Chunbo Luo, Xiaoyang Wang
arXiv Computer Vision
2d ago

CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight

CtrlWAM introduces a controllable world action model that jointly predicts actions (intent) and visual futures (foresight). By executing perturbed actions in a simulator and pairing them with noised visual outcomes, it aligns action predictions with their visual consequences, using warped video–action noise schedules to maintain visual layout responsiveness. The model extends beyond ego‑only control to multiple agent streams, improving action forecasts, video–action agreement, and command following in driving and robotics experiments.

By Chensheng Peng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, Masayoshi Tomizuka, Boris Ivanovic, Marco Pavone, Yuxiao Chen
arXiv Computer Vision
Sep 11

World in World: Explore the World with World Models

World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.

By Chenxi Song, Yanming Yang, Chi Zhang
arXiv Computer Vision
Sep 7

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.

By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
arXiv AI
3d ago

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

arXiv:2609.40219v1 Announce Type: cross Abstract: World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experienc...

By Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan
arXiv Machine Learning
Jul 17

Augmentations for Robust and Efficient Imitation Learning in Streamed Video Games

arXiv:2607. 14200v1 Announce Type: new Abstract: Imitation learning is an appealing way to scale game-playing agents to complex 3D environments by training policies to map visual observations to actions from human demonstrations.

By Somjit Nath, Abdelhak Lemkhenter, Pallavi Choudhury, Chris Lovett, Katja Hofmann, Sergio Valcarcel Macua, Lukas Sch\"afer