arXiv Computer Vision

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Puffin-World is a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation without external offline modules. It jointly models physics, geometry, and appearance as native world states and uses a unified Omni-Camera representation to support diverse tasks and flexible motions. The framework also propagates physical dynamics across future frames, couples appearance and geometry in a single generative process, and scales to complex scenarios with the Puffin-16M dataset of 15 million vision‑language‑camera triplets and 1 million trajectories.

Hugging Face Trending Papers
Sep 3

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Puffin-World is a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation without external offline modules. It jointly models physics, geometry, and appearance as native world states and uses a unified Omni-Camera representation to support diverse tasks and flexible motions. The framework also propagates physical dynamics across future frames, couples appearance and geometry in a single generative process, and scales to complex scenarios with the Puffin-16M dataset of 15 million vision‑language‑camera triplets and 1 million trajectories.

arXiv AI
Jul 2

EgoSim: Egocentric World Simulator for Embodied Interaction Generation

arXiv:2604. 01001v2 Announce Type: replace-cross Abstract: We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation.

By Jinkun Hao, Mingda Jia, Ruiyan Wang, Hongrui Zhu, Jiafei Cao, Xihui Liu, Ran Yi, Lizhuang Ma, Jiangmiao Pang, Xudong Xu
arXiv Computer Vision
Aug 28

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter introduces a two‑stage framework for single‑image world modeling that first generates a global 3D proxy using a Point‑anchored Sparse Structure Flow module, then refines appearance with a Generative Deferred Refiner built on a video diffusion model. The method incorporates Parallel Geometry Injection and Proxy‑Aware Corruption training to integrate the proxy without disrupting the pretrained generative manifold, and it is evaluated on a newly constructed dataset of 115K scenes. Experiments demonstrate that SpatialCrafter outperforms existing approaches, reducing long‑term drift and maintaining consistency under rapid camera motion and extreme viewpoints.

By Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan
Hugging Face Trending Papers
Jul 30

PhiZero: A World Model Built Around Physical Language

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors.

arXiv Computer Vision
Sep 7

TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

TourPhysics is an online framework that builds physics‑grounded visual world models from a single image and a declarative physical configuration. It integrates deterministic simulation with video generation, separating simulator state, geometric evidence, generator controls, and appearance memory to produce consistent observations for exploration and manipulation. The system preserves the input scene, follows prescribed camera and object trajectories more closely than baselines, and reduces appearance drift during long‑horizon revisits.

By Xin Zhang, Yabo Chen, Zixuan Duan, Haibin Huang, Chi Zhang, Feng Xu, Xuelong Li
arXiv Computer Vision
Sep 7

WorldSculpt: Generating Compositional Worlds from Grounded Videos

WorldSculpt presents a method for generating compositional 3D representations of cluttered scenes with hundreds of objects by adapting a single-object 3D generative prior to multi-view observations. The approach, built on Pixal3D with a multi-view conditioning pathway, can generalize to highly occluded scenes without scene-level training. The authors also introduce the UE-MeshyScene benchmark and demonstrate that their method outperforms prior approaches across various evaluation settings, including converting existing 3DGS worlds into compositional mesh scenes.

By Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang