arXiv Computer Vision
Sep 4

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Puffin-World is a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation without external offline modules. It jointly models physics, geometry, and appearance as native world states and uses a unified Omni-Camera representation to support diverse tasks and flexible motions. The framework also propagates physical dynamics across future frames, couples appearance and geometry in a single generative process, and scales to complex scenarios with the Puffin-16M dataset of 15 million vision‑language‑camera triplets and 1 million trajectories.

By Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy
arXiv Computer Vision
Sep 11

Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation

Dream4D is a new framework for generating spatiotemporally coherent 4D content. It uses a two‑stage pipeline: first, few‑shot learning predicts optimal camera trajectories from a single image; second, a pose‑conditioned diffusion process creates geometrically consistent multi‑view sequences that are converted into a persistent 4D representation. The method uniquely combines rich temporal priors from video diffusion models with geometric awareness from reconstruction models, achieving higher quality metrics such as mPSNR and mSSIM compared to existing approaches.

By Xiaoyan Liu, Kangrui Li, Jiaxin Liu, Yuehao Song, Yujie Xing