arXiv AI By Jingyun Liang, Min Wei, Shikai Li, Yizeng Han, Hangjie Yuan, Lei Sun, Weihua Chen, Fan Wang

Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization

Read the original on arXiv AI →

arXiv:2606. 02000v1 Announce Type: cross Abstract: Diffusion models have shown remarkable success in video generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 27

4DStreamCtrl: Interactive Video Generation with Online 4D Control

The paper introduces 4DStreamCtrl, a system that unifies camera motion, object trajectories, and depth into a single 3D point‑track representation, enabling joint control, depth editing, and motion transfer in a single forward pass. By mining in‑the‑wild video for 3D motion supervision and encoding it with a lightweight Geometric Motion Head, the authors train a causal streaming student that can generate arbitrarily long videos in just four denoising steps, achieving 20 FPS on a single high‑end GPU for 480p video. This approach outperforms prior camera‑only, 2D, and offline‑3D methods in motion‑control precision while maintaining temporal coherence over hundreds of frames, thereby enabling interactive 4D‑controllable streaming generation for the first time.

By Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
arXiv Computer Vision
2d ago

4Director: Controlling Video World Models with Rigid 3D Geometry

arXiv:2610.02160v1 Announce Type: new Abstract: Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through i...

By Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
arXiv Computer Vision
Sep 14

VideoTok4D: A 4D-Aware Video Tokenizer for Compact World Representation

VideoTok4D introduces a 4D‑aware video tokenizer that transforms videos into compact tokens representing a dynamic 3D world. It employs spatiotemporal disentanglement to separate static and dynamic content, a track‑aware dynamic attention mechanism for motion consistency across views, and a diffusion prior (Co4DGen) for efficient 4D scene generation. Experiments show state‑of‑the‑art results with up to four orders of magnitude less storage than dense 4D representations, and shorter diffusion sequences for faster generation.

By Xinyi Chen, Hanxin Zhu, Xijun Wang, Xingrui Wang, Sen Liang, Xin Li, Zhibo Chen