AESOP: Asymmetric Human-Camera Generation with Translation-Intensity Control
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 13432v1 Announce Type: cross Abstract: Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control.
arXiv:2609.38683v1 Announce Type: cross Abstract: Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves...
arXiv:2510.24904v2 Announce Type: replace Abstract: Although recent video generative models are getting more capable of following external camera controls, imposed by either text descriptions or came...
Auteur is a language‑driven method that generates human‑centric camera framing for generative video models. It treats shots as framings relative to an actor, encoding shot size, angle, and composition as functions of human pose and motion, and uses a domain‑specific language that converts to standard 6‑DoF camera parameters. A fine‑tuned multimodal large language model acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are interpolated into continuous camera trajectories for video generation.
arXiv:2609.37495v1 Announce Type: new Abstract: Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While exis...
arXiv:2606. 19676v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success in image and video generation and editing.