TKCAM: Text and Keyframe to Camera Trajectory Generation
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2610.11105v1 Announce Type: new Abstract: Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We int...
arXiv:2609.38683v1 Announce Type: cross Abstract: Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves...
Auteur is a language‑driven method that generates human‑centric camera framing for generative video models. It treats shots as framings relative to an actor, encoding shot size, angle, and composition as functions of human pose and motion, and uses a domain‑specific language that converts to standard 6‑DoF camera parameters. A fine‑tuned multimodal large language model acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are interpolated into continuous camera trajectories for video generation.
arXiv:2603.11421v2 Announce Type: replace Abstract: Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Impl...
arXiv:2606. 29531v1 Announce Type: cross Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs.
CameraEditor is a new framework that transforms camera-controlled image editing into a temporal sequence prediction problem. By using video diffusion models, it incorporates a geometric perception module and dynamic reference routing to create precise visual references through dynamic panorama cropping. The method also inserts intermediate transition frames to handle large perspective shifts, maintaining content identity and spatial coherence, and is evaluated on a dataset of 5,760 instances with a benchmark of 462 test cases, achieving state‑of‑the‑art performance.