arXiv Computer Vision

VideoTok4D: A 4D-Aware Video Tokenizer for Compact World Representation

VideoTok4D introduces a 4D‑aware video tokenizer that transforms videos into compact tokens representing a dynamic 3D world. It employs spatiotemporal disentanglement to separate static and dynamic content, a track‑aware dynamic attention mechanism for motion consistency across views, and a diffusion prior (Co4DGen) for efficient 4D scene generation. Experiments show state‑of‑the‑art results with up to four orders of magnitude less storage than dense 4D representations, and shorter diffusion sequences for faster generation.

arXiv Computer Vision
Sep 11

Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation

Dream4D is a new framework for generating spatiotemporally coherent 4D content. It uses a two‑stage pipeline: first, few‑shot learning predicts optimal camera trajectories from a single image; second, a pose‑conditioned diffusion process creates geometrically consistent multi‑view sequences that are converted into a persistent 4D representation. The method uniquely combines rich temporal priors from video diffusion models with geometric awareness from reconstruction models, achieving higher quality metrics such as mPSNR and mSSIM compared to existing approaches.

By Xiaoyan Liu, Kangrui Li, Jiaxin Liu, Yuehao Song, Yujie Xing
arXiv Computer Vision
Aug 25

BulletTime: Decoupled Control of Time and Camera Pose for Video Generation

BulletTime presents a 4D‑controllable video diffusion framework that separates scene dynamics from camera pose, allowing precise manipulation of both temporal and spatial aspects of generated videos. The model conditions on continuous world‑time sequences and camera trajectories, integrating them via a 4D positional encoding in the attention layer and adaptive normalizations for feature modulation. A specially curated dataset with independently parameterized temporal and camera variations is used for training, and the resulting system demonstrates robust real‑world 4D control while maintaining high generation quality and surpassing prior methods in controllability.

By Yiming Wang, Qihang Zhang, Shengqu Cai, Tong Wu, Jan Ackermann, Zhengfei Kuang, Yang Zheng, Frano Raji\v{c}, Siyu Tang, Gordon Wetzstein
arXiv Computer Vision
Sep 21

4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors

The paper introduces 4DGS-Fixer, an iterative refinement framework that uses a video diffusion model to enhance sparse-view 4D Gaussian Splatting for dynamic scene synthesis. It first fuses multi-view depth maps into dense point clouds for better geometric initialization, then applies a pretrained video restoration model to refine rendered sequences, providing pseudo-supervision for further refinement. Experiments on a benchmark dataset show the method outperforms existing baselines, achieving nearly a 2 dB PSNR improvement.

By Haitao Huang, Shenghao Zhao, Boyuan Tian, Shin-Fang Chng, Songlin Yang, Sheila Lim, Huangying Zhan, Yi Xu, Anyi Rao, Frank Guan
Hugging Face Trending Papers
Jul 21

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding.

arXiv Computer Vision
2d ago

4Director: Controlling Video World Models with Rigid 3D Geometry

arXiv:2610.02160v1 Announce Type: new Abstract: Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through i...

By Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
arXiv Computer Vision
Sep 1

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

arXiv:2608.18734v2 Announce Type: replace Abstract: 4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vis...

By Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe, Hasitha Gallella, Dulanga Weerakoon, Vigneshwaran Subbaraju, Ranga Rodrigo
arXiv Computer Vision
Sep 15

Stereo4DWalker: Learning 4D-aware Embodied Urban Navigation from Internet Stereo Videos

Stereo4DWalker is a 4D-aware embodied navigation model that uses stereo video inputs to explicitly construct structured representations of geometry and motion. These 4D structures are incorporated into a navigation transformer via 4D-conditioned attention layers, enabling the agent to learn robust urban navigation. The authors also curate a large-scale stereo navigation dataset with automatically annotated actions from Internet stereo videos, and demonstrate that Stereo4DWalker outperforms state‑of‑the‑art methods while requiring only 1.5% of the training data.

By Wentao Zhou, Xuweiyi Chen, Vignesh Rajagopal, Jeffrey Chen, Rohan Chandra, Zezhou Cheng