State of open video generation models in Diffusers
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes slow generation.
ReaDiT Guidance is a lightweight framework that controls Diffusion Transformer (DiT) models by leveraging internal feature representations from a single DiT block. It steers image and video generation toward spatial targets such as depth, pose, or edge maps, and extends naturally to video generation for camera and motion control. Experiments show competitive or improved results compared to existing feature‑based and adapter‑based methods while using fewer parameters.
We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios.