The paper introduces 4DStreamCtrl, a system that unifies camera motion, object trajectories, and depth into a single 3D point‑track representation, enabling joint control, depth editing, and motion transfer in a single forward pass. By mining in‑the‑wild video for 3D motion supervision and encoding it with a lightweight Geometric Motion Head, the authors train a causal streaming student that can generate arbitrarily long videos in just four denoising steps, achieving 20 FPS on a single high‑end GPU for 480p video. This approach outperforms prior camera‑only, 2D, and offline‑3D methods in motion‑control precision while maintaining temporal coherence over hundreds of frames, thereby enabling interactive 4D‑controllable streaming generation for the first time.
By Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
arXiv:2610.02160v1 Announce Type: new
Abstract: Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through i...
By Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
arXiv:2605. 13838v3 Announce Type: replace-cross Abstract: Video-guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets.
By Zijie Wu, Lixin Xu, Puhua Jiang, Sicong Liu, Chunchao Guo, Xiang Bai
arXiv:2607. 17097v2 Announce Type: replace Abstract: Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI.
By Lingwei Dang, Juntong Li, Zonghan Li, Hongwen Zhang, Liang An, Wei Min, Yebin Liu, Qingyao Wu
VideoTok4D introduces a 4D‑aware video tokenizer that transforms videos into compact tokens representing a dynamic 3D world. It employs spatiotemporal disentanglement to separate static and dynamic content, a track‑aware dynamic attention mechanism for motion consistency across views, and a diffusion prior (Co4DGen) for efficient 4D scene generation. Experiments show state‑of‑the‑art results with up to four orders of magnitude less storage than dense 4D representations, and shorter diffusion sequences for faster generation.
By Xinyi Chen, Hanxin Zhu, Xijun Wang, Xingrui Wang, Sen Liang, Xin Li, Zhibo Chen
arXiv:2608.23549v1 Announce Type: new
Abstract: Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces...
By Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan