DiVid: Diagnosing Dimension-Specific Diversity Collapse in Video Generation Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 09024v1 Announce Type: cross Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models.
arXiv:2606. 10620v1 Announce Type: cross Abstract: Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly understood.
WanPE is a 397‑B parameter prompt‑enhancement model that learns director‑level cinematic planning from 1.05 M real‑world videos. It generates shot‑level cinematic plans through video‑grounded reverse construction and uses Semantic‑Consistency GRPO (SC‑GRPO) to maintain user intent across shots and time. In evaluations, WanPE improves human preference over raw prompts by up to 50.86 points for 30‑second videos and outperforms commercial offerings for shorter durations.
Auteur is a language‑driven method that generates human‑centric camera framing for generative video models. It treats shots as framings relative to an actor, encoding shot size, angle, and composition as functions of human pose and motion, and uses a domain‑specific language that converts to standard 6‑DoF camera parameters. A fine‑tuned multimodal large language model acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are interpolated into continuous camera trajectories for video generation.
arXiv:2608.23383v1 Announce Type: new Abstract: Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, foll...
arXiv:2609.37407v1 Announce Type: new Abstract: While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a ma...