arXiv AI By Samuel Garcin, Thomas Walker, Steven McDonagh, Tim Pearce, Hakan Bilen, Tianyu He, Kaixin Wang, Jiang Bian

Beyond Pixel Histories: World Models with Persistent 3D State

Read the original on arXiv AI →

arXiv:2603. 03482v2 Announce Type: replace-cross Abstract: Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 28

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter introduces a two‑stage framework for single‑image world modeling that first generates a global 3D proxy using a Point‑anchored Sparse Structure Flow module, then refines appearance with a Generative Deferred Refiner built on a video diffusion model. The method incorporates Parallel Geometry Injection and Proxy‑Aware Corruption training to integrate the proxy without disrupting the pretrained generative manifold, and it is evaluated on a newly constructed dataset of 115K scenes. Experiments demonstrate that SpatialCrafter outperforms existing approaches, reducing long‑term drift and maintaining consistency under rapid camera motion and extreme viewpoints.

By Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan
arXiv Computer Vision
Sep 1

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

arXiv:2608.29910v1 Announce Type: new Abstract: Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling appli...

By Runjia Qian, Zile Wang, Jihai Zhang, Kai Zou, Wei Yu, Jiaxing Li, Zexiang Liu, Yaokun Li, Fei Kang, Kaichen Huang, Mengyin An, Haobo Zhang, Biao Jiang, Jiahua Wang, Haofeng Sun, Yang Liu, Yangguang Li
arXiv Computer Vision
Sep 10

Programmable World Model

arXiv:2609.10540v1 Announce Type: new Abstract: Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent w...

By Zheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang, Ruihan Yu, Muyao Niu, Siqi Yang, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
arXiv Computer Vision
Sep 22

GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation

The paper introduces the Geometry‑Native Autoencoder (GAE), a compact latent space that can be decoded into appearance, depth, camera parameters, and point maps, enabling 3D‑consistent world generation. By reparameterizing a geometry foundation model’s features, GAE replaces traditional appearance‑centric latents and improves visual quality and 3D coherence, achieving significant reductions in FVD and camera‑trajectory error on benchmark datasets. The work demonstrates that a geometry‑native latent space can serve as a shared interface between perception and generation models.

By Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu
arXiv Computer Vision
Sep 4

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Puffin-World is a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation without external offline modules. It jointly models physics, geometry, and appearance as native world states and uses a unified Omni-Camera representation to support diverse tasks and flexible motions. The framework also propagates physical dynamics across future frames, couples appearance and geometry in a single generative process, and scales to complex scenarios with the Puffin-16M dataset of 15 million vision‑language‑camera triplets and 1 million trajectories.

By Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy