arXiv AI

GraphVid: Interactive Graph-Controllable Video Generation

arXiv:2607. 21580v1 Announce Type: cross Abstract: Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement.

arXiv Computer Vision
6d ago

Bootstrapping Video Interaction Generation with Synthetic State Transitions

The paper presents a framework for creating a scalable synthetic dataset of controllable video interactions by generating explicit start and end state images using image editing models. It introduces State‑Guided Sampling (SGS) to produce seamless videos anchored on these states, reducing artifacts seen in naive conditional generation. An automated evaluation system aligned with human judgments is also developed, and experiments demonstrate that fine‑tuning a base model on this dataset markedly improves its ability to generate plausible interactions.

By Jiho Jang, Jinyoung Kim, Nojun Kwak, Kyungjune Kim
arXiv Computer Vision
1d ago

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

ALIVE is a new framework for first‑frame‑guided video editing that enables inserted objects to participate in coherent interactions with the source video, such as being picked up or manipulated. The authors curated 35,800 editing pairs from 3D‑rendered, model‑generated, and real‑world videos, and trained a vision‑language model to predict interaction guidance from the edited first frame and an object name. ALIVE outperforms the strongest baseline by 43.9% on overall performance and 4.4% on a general video object insertion benchmark, with VLM‑predicted guidance further improving interaction fidelity.

By Zhenghong Zhou, Zhe Lin, Jiebo Luo, Yuqian Zhou
arXiv Computer Vision
Aug 27

4DStreamCtrl: Interactive Video Generation with Online 4D Control

The paper introduces 4DStreamCtrl, a system that unifies camera motion, object trajectories, and depth into a single 3D point‑track representation, enabling joint control, depth editing, and motion transfer in a single forward pass. By mining in‑the‑wild video for 3D motion supervision and encoding it with a lightweight Geometric Motion Head, the authors train a causal streaming student that can generate arbitrarily long videos in just four denoising steps, achieving 20 FPS on a single high‑end GPU for 480p video. This approach outperforms prior camera‑only, 2D, and offline‑3D methods in motion‑control precision while maintaining temporal coherence over hundreds of frames, thereby enabling interactive 4D‑controllable streaming generation for the first time.

By Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
arXiv AI
6d ago

Generative Cinematographer: Composing Camera and Object Motion in 3D

Generative Cinematographer (GenCine) lifts a single image into an editable 3D scene scaffold, allowing artists to jointly author camera paths and foreground motion using local 3D motion handles. The system projects these controls into guidance maps that encode handle positions and colors across frames, enabling a pretrained video model to follow camera-relative, piecewise-rigid motion without physics simulation. Training recovers controls from real videos and synthetic ground-truth geometry, and experiments demonstrate consistent camera-relative motion, improved geometric consistency, and strong controllability across diverse real-world scenes.

By Jiahan Zhang, Chaohao Yang, Namitha Guruprasad, Vivekjyoti Banerjee, Trong-Tung Nguyen, Alan Yuille, Anand Bhattad
arXiv Computer Vision
Sep 25

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

WanPE is a 397‑B parameter prompt‑enhancement model that learns director‑level cinematic planning from 1.05 M real‑world videos. It generates shot‑level cinematic plans through video‑grounded reverse construction and uses Semantic‑Consistency GRPO (SC‑GRPO) to maintain user intent across shots and time. In evaluations, WanPE improves human preference over raw prompts by up to 50.86 points for 30‑second videos and outperforms commercial offerings for shorter durations.

By Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
Hugging Face Trending Papers
Jun 17

LooseControlVideo: Directorial Video Control using Spatial Blocking

Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and temporal dynamics are often entangled. While existing depth-conditioned models achieve good structural fidelity, they necessitate dense, frame-accurate guidance that is labor-intensive to author for dynamic events involving deformable objects.

arXiv Computer Vision
Sep 22

VideoGen-Agent: Reinforcing Video Generation Agents

VideoGen-Agent is a multimodal agent that uses multitask agentic reinforcement learning to coordinate external tools for video generation. It learns to augment, generate, and verify videos through multi‑turn interactions, guided by prompts and intermediate observations. On the new VABench benchmark, the agent improves base text‑to‑video performance by 19.1 points, and further upgrades to generation tools raise the score to 86.1, with human raters favoring the upgraded configuration in 84.3% of comparisons.

By Binxu Li, Haoyi Duan, Yuhui Zhang, Yaohui Zhang, Zihao Lin, Kaituo Feng, Suozhi Huang, Xiangyi Li, Yu Li, Chunyuan Li, Shilong Liu, Mengdi Wang