arXiv Computer Vision

Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models

arXiv Computer Vision
Sep 23

Code Plans, Diffusion Renders: Open-Ended Generative World Modeling

The paper introduces CoDeR, a new paradigm for world modeling that explicitly builds an executable world using code rather than relying solely on visual observations. CoDeR translates high‑level concepts into structured world rules, executable dynamics, and perceptual observations through five complementary roles, enabling long‑term memory, open‑ended interactions, autonomous world evolution, and persistent multi‑agent dynamics. Experiments show that this framework extends the capabilities of existing world models and achieves state‑of‑the‑art performance across multiple evaluation settings.

By Zixun Fang, Yawen Shao, Kai Zhu, Jie Xiao, Shihan Chen, Yu Liu, Xueyang Fu, Yang Cao, Wei Zhai, Zheng-Jun Zha
arXiv Computer Vision
Sep 11

World in World: Explore the World with World Models

World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.

By Chenxi Song, Yanming Yang, Chi Zhang
arXiv AI
Aug 28

GameWAM: A World Action Model for Video Games

GameWAM is the first World-Action Model designed for native closed-loop gameplay and GUI control in modern video games. It jointly generates future visual observations and executable keyboard-mouse trajectories using parallel visual and action generative processes, block-causal conditioning, and flow matching. The model predicts gameplay/GUI mode at each step, handles heterogeneous native controls, and employs block-cycle control for long-horizon interaction, achieving competitive task success with fewer native actions than prior agents.

By Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li
arXiv Computation and Language
Aug 27

Code World Model: Coding Agent as World Brain

The paper introduces Code World Model, a framework that decouples world evolution from visual rendering by using a coding agent as a world brain. The agent reasons about events, generates executable code to maintain persistent state, and a proxy representation links this state to a video model for high‑fidelity visual output. Experiments with MiniMax‑H3 show that the system can follow proxy‑based spatiotemporal specifications while preserving rich visual dynamics, illustrating a new approach to open‑ended world modeling.

By Yiwen Chen, Guosheng Lin, Chi Zhang
arXiv Computer Vision
Sep 10

Programmable World Model

arXiv:2609.10540v1 Announce Type: new Abstract: Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent w...

By Zheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang, Ruihan Yu, Muyao Niu, Siqi Yang, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
arXiv AI
Jul 22

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.

By AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
arXiv Computer Vision
Sep 7

TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

TourPhysics is an online framework that builds physics‑grounded visual world models from a single image and a declarative physical configuration. It integrates deterministic simulation with video generation, separating simulator state, geometric evidence, generator controls, and appearance memory to produce consistent observations for exploration and manipulation. The system preserves the input scene, follows prescribed camera and object trajectories more closely than baselines, and reduces appearance drift during long‑horizon revisits.

By Xin Zhang, Yabo Chen, Zixuan Duan, Haibin Huang, Chi Zhang, Feng Xu, Xuelong Li