arXiv Computer Vision
Sep 25

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation

ViRDM is a new post‑training method for few‑step causal video generation that eliminates the need for a large teacher model and an online critic. By applying representation distribution matching (RDM) with a precomputed target distribution, a lightweight VAE decoder, and staged vector–Jacobian products, ViRDM overcomes memory, optimization, and temporal dynamics challenges. The approach reduces GPU memory usage and training time, achieving state‑of‑the‑art VBench performance with only 20 generator updates and 16 A100 GPU‑hours.

By Zichong Meng, Chongjian Ge, Chun-Hao P. Huang, Yang Zhou, Huaizu Jiang
arXiv AI
6d ago

DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

DyMD introduces a Distribution Matching Distillation framework that adapts teacher supervision and critic fitting to preserve interaction dynamics in few-step video generation. By employing temporal affinity–conditioned re‑noise sampling and dynamics‑guided fake‑score tracking, DyMD balances motion recovery with visual quality. The method distills a 14B teacher into a 1.3B student that achieves significant gains on embodied‑video benchmarks and downstream action planning tasks.

By Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu