arXiv Computer Vision

Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware

Waypoint‑1.5 is a real‑time diffusion world model designed for interactive video generation on consumer‑grade hardware. It is pre‑trained on 100,000 hours of control‑aligned video game footage and can generate playable video conditioned on full keyboard and mouse input. The system offers two resolution variants, distinguishes rendered FPS, latent FPS, and control rate, and includes a detailed data pipeline, architecture, training methodology, and runtime system.

arXiv Computer Vision
2d ago

Matrix-game 2.0: An open-source, real-time, and streaming interactive world model

arXiv:2508.13009v5 Announce Type: replace Abstract: Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynami...

By Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Size Wu, Wei Li, Xuchen Song, Yang Liu, Yangguang Li, Yahui Zhou
arXiv Computer Vision
Sep 10

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

arXiv:2607.26694v3 Announce Type: replace Abstract: We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generat...

By Xiangbo Gao, Siyuan Yang, Ping He, Mingyang Wu, Yuheng Wu, Yushen Zuo, Jiongze Yu, Ryan Cui, Hongyuan Hua, Devin Ma, Xiao Jin, Yubo Ruan, Qing Yin, Jie Yang, Zhengzhong Tu
arXiv AI
Jul 22

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.

By AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
arXiv Computer Vision
Aug 28

Magpie: Real-Time World Renderer for Interactive Games

Magpie is a real‑time generative world‑rendering system designed for interactive games. It decouples gameplay execution from visual generation, allowing designers to define scenes and rules in a game engine while a separate render server produces visual output from the engine’s white‑box frames. This approach preserves gameplay designability and reproducibility, and reduces early prototype dependence on finished visual assets.

By Xiaoyu Zhan, Xinyu Wang, Xiaohong Zhang, Huanjie Zhu, Tengjiao Sun, Pengcheng Fang, Jiaxing Yu, Yanwen Guo, Dongjie Fu
arXiv AI
Jul 22

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

arXiv:2607. 19191v1 Announce Type: cross Abstract: We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics.

By Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
arXiv AI
Jul 23

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation

arXiv:2607. 20174v1 Announce Type: cross Abstract: Existing human--object interaction (HOI) video generation methods are largely limited to offline short-video generation with complex driving conditions, making them unsuitable for real-time interactive applications.

By Zejing Rao, Haoxian Zhang, Xiaoqiang Liu, Yiping Meng, Guoxin Zhang, Pengfei Wan, Fan Tang, Tong-Yee Lee
arXiv Machine Learning
Jul 7

Vidu S1: A Real-Time Interactive Video Generation Model

arXiv:2607. 03118v1 Announce Type: cross Abstract: We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters.

By Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Yang Luo, Yuji Wang, Dechuang Chen, Jungang Li, Chengyang Ye, Marco Chen, Hongzhou Zhu, Min Zhao, Yuxuan Jiang, Zhengkun Huang, Chendong Xiang, Kaiwen Zheng, Haoxu Wang, Xiaohang Wang, Qi Jia, Xin Chen, Yimin Chen, Youhe Jiang, Fangcheng Fu, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu
arXiv Computer Vision
Sep 4

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

The paper introduces a large‑scale synthetic data pipeline built on Unreal Engine to generate action‑conditioned, multi‑view video for training world models. The system operates in two stages: real‑time physics simulation records trajectories, then offline rendering produces high‑quality video. It includes a distributed production framework with task partitioning, automated filtering, and a 25‑server cluster, yielding over 2,600 hours of 1080p and 6,000 hours of 720p video from 429 levels and 40 characters.

By Haoyu Wang, Songchun Zhang, Haoran Li, Haoyang Huang, Zeyue Xue, Nan Duan