A Study of the Design Space of Motion and Disocclusion Control in Video Generation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and temporal dynamics are often entangled. While existing depth-conditioned models achieve good structural fidelity, they necessitate dense, frame-accurate guidance that is labor-intensive to author for dynamic events involving deformable objects.
arXiv:2609.36940v1 Announce Type: new Abstract: Accurate dynamic scene reconstruction is important for robotic perception, where temporally consistent representations of dynamic environments are esse...
arXiv:2610.02180v1 Announce Type: cross Abstract: Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambig...
arXiv:2606. 02000v1 Announce Type: cross Abstract: Diffusion models have shown remarkable success in video generation.
arXiv:2610.02160v1 Announce Type: new Abstract: Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through i...
arXiv:2609.37495v1 Announce Type: new Abstract: Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While exis...