arXiv Computer Vision

StableWorld: Towards Stable and Consistent Long Interactive Video Generation

arXiv AI
Jul 22

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.

By AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
arXiv Computer Vision
Sep 10

Programmable World Model

arXiv:2609.10540v1 Announce Type: new Abstract: Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent w...

By Zheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang, Ruihan Yu, Muyao Niu, Siqi Yang, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
arXiv Computer Vision
2d ago

Matrix-game 2.0: An open-source, real-time, and streaming interactive world model

arXiv:2508.13009v5 Announce Type: replace Abstract: Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynami...

By Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Size Wu, Wei Li, Xuchen Song, Yang Liu, Yangguang Li, Yahui Zhou
arXiv Computer Vision
Sep 1

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

arXiv:2608.29910v1 Announce Type: new Abstract: Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling appli...

By Runjia Qian, Zile Wang, Jihai Zhang, Kai Zou, Wei Yu, Jiaxing Li, Zexiang Liu, Yaokun Li, Fei Kang, Kaichen Huang, Mengyin An, Haobo Zhang, Biao Jiang, Jiahua Wang, Haofeng Sun, Yang Liu, Yangguang Li
arXiv Computer Vision
Sep 22

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

WorldCrafter is a video world model that introduces a camera‑queryable implicit 3D‑aware memory to improve long‑horizon consistency and viewpoint control. The model compresses multi‑view evidence into a limited token budget shaped by the requested viewpoint, integrating historical observations via a memory encoder and pose‑conditioned readout before denoising. Experiments on static and dynamic scenes demonstrate significant gains in consistency and camera‑control accuracy while maintaining visual quality during minute‑scale exploration.

By Wangbo Yu, Kunhao Liu, Wenbo Hu, Shenghai Yuan, Chaoran Feng, Haiyang Zhou, Yukun Huang, Yiran Wang, Wang Zhao, Yingmin Luo, Ying Shan
arXiv Computer Vision
2d ago

MUGEN: Interactive Panoramic World Exploration via Camera Control

MUGEN is a large‑scale real‑world dataset of over 1,300 hours of 4K panoramic videos with rich semantic and geometric annotations, designed to support interactive 360° world exploration. Wan360 is a camera‑controllable panoramic video generation model built on MUGEN, featuring ERP‑aware components (periodic longitude RoPE, ERP‑aware padding, random roll yaw) and a panoramic Plücker embedding for camera motion. Together, they address gaps in data and model support for immersive, temporally coherent 360° video generation along user‑specified camera trajectories.

By Jiaming Tan, Zhen Li, Shuwei Shi, Minggui Teng, Siqi Yang, Yuwei Wu, Bo Zheng, Chuanhao Li, Kaipeng Zhang
arXiv Computer Vision
Aug 25

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

arXiv:2608.23383v1 Announce Type: new Abstract: Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, foll...

By Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
arXiv Computer Vision
Sep 11

World in World: Explore the World with World Models

World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.

By Chenxi Song, Yanming Yang, Chi Zhang