arXiv Computer Vision

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training

arXiv AI
Jul 22

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.

By AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
arXiv Computer Vision
Aug 25

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

arXiv:2607.14935v2 Announce Type: replace Abstract: Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world application...

By Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
arXiv Computer Vision
Aug 25

WorldMind: Decoupled Game World Model for State-Aware NPC Behavior

WorldMind is a decoupled framework for state-aware NPC behavior in game world models, separating interactive world modeling into four layers: Understanding, Decision, Control, and Generation. It constructs a compact state from generated frames, reasons over it to plan NPC actions, translates actions into temporally aligned conditions, and synthesizes visual outcomes. Experiments on the newly introduced BOSS-140K dataset show that WorldMind achieves more tactically appropriate and coherent NPC behavior than baseline models in about 70% of pairwise comparisons.

By Zhiyang Deng, Boran Zhang, Danze Chen, Yeying Jin
arXiv Computation and Language
Aug 27

Code World Model: Coding Agent as World Brain

The paper introduces Code World Model, a framework that decouples world evolution from visual rendering by using a coding agent as a world brain. The agent reasons about events, generates executable code to maintain persistent state, and a proxy representation links this state to a video model for high‑fidelity visual output. Experiments with MiniMax‑H3 show that the system can follow proxy‑based spatiotemporal specifications while preserving rich visual dynamics, illustrating a new approach to open‑ended world modeling.

By Yiwen Chen, Guosheng Lin, Chi Zhang
arXiv AI
Aug 11

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

arXiv:2607. 27380v2 Announce Type: replace-cross Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt.

By Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
arXiv AI
6d ago

GameWAM: A World Action Model for Video Games

GameWAM is the first World-Action Model designed for native closed-loop gameplay and GUI control in modern video games. It jointly generates future visual observations and executable keyboard-mouse trajectories using parallel visual and action generative processes, block-causal conditioning, and flow matching. The model predicts gameplay/GUI mode at each step, handles heterogeneous native controls, and employs block-cycle control for long-horizon interaction, achieving competitive task success with fewer native actions than prior agents.

By Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li
arXiv Computer Vision
23h ago

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

SolarWM is an open foundation for building interactive video world models, offering a reconfigurable multi‑source data engine that unifies 1.43 million clips from 10 datasets into a consistent, frame‑aligned format. It provides a backbone‑native adaptation framework that preserves native representations of models ranging from 5 B to 33 B parameters, and a three‑stage training recipe combining bidirectional adaptation, teacher‑forced autoregressive initialization, and distribution‑matching distillation. The resulting causal models can interact in real‑time over rollouts from minutes to hours, trained only on 5‑second sequences, and the project releases data, pipeline, recipes, weights, and framework for reproducible research.

By Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang
arXiv Computer Vision
Aug 21

ID-V2V: Identity-Preserving Video Restylization

arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.

By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu