arXiv Computer Vision

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

arXiv AI
Jul 22

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.

By AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
arXiv Computer Vision
Sep 4

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

The paper introduces a large‑scale synthetic data pipeline built on Unreal Engine to generate action‑conditioned, multi‑view video for training world models. The system operates in two stages: real‑time physics simulation records trajectories, then offline rendering produces high‑quality video. It includes a distributed production framework with task partitioning, automated filtering, and a 25‑server cluster, yielding over 2,600 hours of 1080p and 6,000 hours of 720p video from 429 levels and 40 characters.

By Haoyu Wang, Songchun Zhang, Haoran Li, Haoyang Huang, Zeyue Xue, Nan Duan
arXiv Computer Vision
Sep 15

DiVA: Enabling Interactive Digital Life Simulation via Video Models

arXiv:2609.13830v1 Announce Type: new Abstract: We present DiVA, a deeply interactive digital life simulator pioneering a new paradigm for long-term, open-ended interactive experiences within digital...

By Cheng Chen, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Wen Wang, Yihao Meng, Hanlin Wang, Yixuan Li, Jiacheng Wei, Zhenshan Tan, Yanhong Zeng, Yujun Shen, Guosheng Lin, Fayao Liu
arXiv Computer Vision
Aug 25

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

arXiv:2608.23383v1 Announce Type: new Abstract: Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, foll...

By Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
arXiv Machine Learning
Sep 11

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Vidu S2 is a system that includes Vidu S2-Avatar, a real‑time interactive digital‑character model, and Vidu S2-Editing, a real‑time video editing model. It enables real‑time 720p video generation with dynamic references and improved instruction following, such as dancing, and allows real‑time editing of video streams for style rendering, clothing replacement, character replacement, and background replacement. Experiments show Vidu S2 outperforms all baselines, and a playable online demo is available at https://vidu.com/vidu-stream.

By Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li, Dechuang Chen, Ming Lin, Jingjiang Zhou, Haopeng Jin, Qi Jia, Xiaohang Wang, Yaole Wang, Zhanqiang Zhang, Ran Li, Zhengkun Huang, Shuyue Xiong, Yuji Wang, Zikun Dai, Hui He, Yang Luo, Mang Ning, Weiqi Feng, Chengyang Ye, Xinyue Lin, Min Zhao, Hongzhou Zhu, Hengkai Tan, Zeyuan Wang, Chendong Xiang, Kaiwen Zheng, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu
arXiv Computer Vision
2d ago

Matrix-game 2.0: An open-source, real-time, and streaming interactive world model

arXiv:2508.13009v5 Announce Type: replace Abstract: Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynami...

By Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Size Wu, Wei Li, Xuchen Song, Yang Liu, Yangguang Li, Yahui Zhou
arXiv AI
Jul 23

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation

arXiv:2607. 20174v1 Announce Type: cross Abstract: Existing human--object interaction (HOI) video generation methods are largely limited to offline short-video generation with complex driving conditions, making them unsuitable for real-time interactive applications.

By Zejing Rao, Haoxian Zhang, Xiaoqiang Liu, Yiping Meng, Guoxin Zhang, Pengfei Wan, Fan Tang, Tong-Yee Lee
arXiv AI
Aug 11

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

arXiv:2607. 27380v2 Announce Type: replace-cross Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt.

By Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng