Bootstrapping Video Interaction Generation with Synthetic State Transitions
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.38172v1 Announce Type: cross Abstract: Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. How...
arXiv:2607. 21580v1 Announce Type: cross Abstract: Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement.
arXiv:2510.24904v2 Announce Type: replace Abstract: Although recent video generative models are getting more capable of following external camera controls, imposed by either text descriptions or came...
arXiv:2604. 06010v2 Announce Type: replace Abstract: Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed.
arXiv:2609.37004v1 Announce Type: new Abstract: We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While ex...
arXiv:2608. 05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.