The paper introduces a promptable localized motion representation that generates persistent embeddings for user-specified regions in a video, without cropping or masking the input. By conditioning motion encoding directly on spatial masks while processing the full video, the method produces temporally consistent, region-addressable embeddings that capture local dynamics while preserving global context. These embeddings enable object-level motion transfer for dynamic scene composition and improve localized action classification in multi-actor videos, outperforming global representations that rely on cropping or post-hoc masking.
By Frank Fundel, Malek Ben Alaya, Thomas Ressler-Antal, Stefan Andreas Baumann, Bj\"orn Ommer
arXiv:2508.13009v5 Announce Type: replace
Abstract: Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynami...
By Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Size Wu, Wei Li, Xuchen Song, Yang Liu, Yangguang Li, Yahui Zhou
GameWAM is the first World-Action Model designed for native closed-loop gameplay and GUI control in modern video games. It jointly generates future visual observations and executable keyboard-mouse trajectories using parallel visual and action generative processes, block-causal conditioning, and flow matching. The model predicts gameplay/GUI mode at each step, handles heterogeneous native controls, and employs block-cycle control for long-horizon interaction, achieving competitive task success with fewer native actions than prior agents.
By Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li
arXiv:2607. 05352v1 Announce Type: cross Abstract: We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions.
By Anthony Hu, V\'aclav Volhejn, Adrien Ramanana Rahary, Chris Mulder, Aditya Makkar, Am\'elie Royer, Manu Orsini, Alyx Liao, Adam Jelley, Eloi Alonso, Florian Laurent, Fredrik Nor\'en, James Swingos, Jan H\"unermann, Kent Rollins, Lucas Hosseini, Matthieu Le Cauchois, Maxim Peter, Pim de Witte, Tim Brown, Vincent Micheli, Moritz B\"ohle, Gabriel de Marmiesse, Viktoriia Sharmanska, Lucia Specia, Michael Black, Patrick P\'erez
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings.
arXiv:2606. 02753v1 Announce Type: cross Abstract: Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective.
By Teng Hu, Mingchun Lu, Yating Wang, Jiangning Zhang, Jinkun Hao, Ye Pan, Ran Yi, Lizhuang Ma, Dacheng Tao
WorldMind is a decoupled framework for state-aware NPC behavior in game world models, separating interactive world modeling into four layers: Understanding, Decision, Control, and Generation. It constructs a compact state from generated frames, reasons over it to plan NPC actions, translates actions into temporally aligned conditions, and synthesizes visual outcomes. Experiments on the newly introduced BOSS-140K dataset show that WorldMind achieves more tactically appropriate and coherent NPC behavior than baseline models in about 70% of pairwise comparisons.
By Zhiyang Deng, Boran Zhang, Danze Chen, Yeying Jin
CounterVid introduces a scalable counterfactual video generation framework that creates videos differing only in actions or temporal structure while keeping scene context intact. The approach uses multimodal LLMs for action proposals and diffusion models for editing, producing a synthetic dataset of ~26k preference pairs for action recognition and sequence ordering. With the MixDPO optimization method, the authors demonstrate significant improvements in action recognition and temporal ordering on Qwen2.5‑VL and InternVL3 backbones, while maintaining overall video understanding.
By Tobia Poppi, Burak Uzkent, Amanmeet Garg, Lucas Porto, Garin Kessler, Yezhou Yang, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara, Florian Schiffers
CLAP is a cross-embodiment framework for action‑conditioned video generation that can be trained on diverse internet‑scale videos from both humans and robots. It reconciles different action spaces—end‑effector poses, language instructions, and latent actions—using a curriculum that first learns physics priors from unlabeled video and then grounds them in real‑world action spaces for zero‑shot deployment. The resulting models match or exceed state‑of‑the‑art single‑embodiment models in challenging environments and support few‑shot adaptation across a wide range of robot morphologies.
By Kechen Liu, Ola Shorinwa
arXiv:2607. 03964v1 Announce Type: cross Abstract: World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulators in which agents can perceive, act, forecast, and acquire scalable experience.
By Jianjie Fang, Yongyan Xu, Ziyou Wang, Chen Gao, Yuchao Huang, Zhaolu Wang, Rongze Tang, Mingyuan Jia, Baining Zhao, Weichen Zhang, Xin Zhang, Haisheng Su, Yu Shang, Wei Wu, Xinlei Chen, Yong Li
arXiv:2603. 02697v2 Announce Type: replace-cross Abstract: This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works that lack support for unified shared world construction with multi-agent interaction.
By Jiayi Zhu, Jianing Zhang, Yiying Yang, Wei Cheng, Xiaoyun Yuan
arXiv:2610.01614v1 Announce Type: new
Abstract: Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended gene...
By Xindi Yang, Baolu Li, Liam Lee, Zhenfei Yin, Songxin Zhang, Zhuoyang Song, Xu Jia, Jianfei Cai, Tien-Tsin Wong, Bingyi Jing, Mengyue Yang