arXiv Computer Vision By Jingzhong Lin, Zhanke Wang, Heng Li, Wenxiang Liu, Zhao Zhang, Kecheng Tang, Dongdong Xiang, Changbo Wang, Di Kang, Chunchao Guo, Linchao Bao, Gaoqi He

AESOP: Asymmetric Human-Camera Generation with Translation-Intensity Control

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

arXiv Computer Vision
6d ago

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation

Auteur is a language‑driven method that generates human‑centric camera framing for generative video models. It treats shots as framings relative to an actor, encoding shot size, angle, and composition as functions of human pose and motion, and uses a domain‑specific language that converts to standard 6‑DoF camera parameters. A fine‑tuned multimodal large language model acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are interpolated into continuous camera trajectories for video generation.

By Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra, Xuelin Chen, Erkut Erdem, Aykut Erdem, Duygu Ceylan