Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.
The paper introduces VoRTeC, a video compression framework that leverages a foundational flow model to encode latent video representations compactly and predict their positions along flow trajectories. By integrating multi‑scale priors and avoiding access to flow‑matching network parameters, VoRTeC achieves one‑step decoding with high perceptual fidelity, while maintaining temporal consistency through tail‑frame reuse and prior caching. Experiments show a 58% reduction in bit consumption compared to prior diffusion‑based methods and a decoding speed increase ranging from 3 to 197 times, reaching 13 FPS at 720p and 32 FPS at 480p.
By Yichong Xia, Qinhong Wu, Qinhong Wu, Jinpeng Wang, Zeyuan Chen, Haoqian Wang
arXiv:2512.07480v2 Announce Type: replace
Abstract: While traditional and neural video codecs (NVCs) have achieved remarkable rate-distortion performance, improving perceptual quality at low bitrates...
By Naifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li, Zihan Zheng, Yuan Zhang, Yan Lu
arXiv:2603.17546v2 Announce Type: replace
Abstract: Perceptual video compression leverages generative priors to reconstruct realistic textures and motions at low bitrates. However, existing perceptua...
By Daowen Li, Ruixiao Dong, Kai Li, Ying Chen, Ding Ding, Li Li
The paper introduces a zero‑shot subject‑driven video generation framework that eliminates the need for per‑subject tuning and large subject‑video datasets. It achieves this by separating identity injection—learned from subject‑image pairs—and motion‑awareness preservation—maintained with a small set of arbitrary videos, and optimizes both with stochastic switching and dropout techniques. Using CogVideoX‑5B, the method adapts a single model with only 200K subject‑image pairs and 4,000 arbitrary videos in 288 A100 GPU hours, representing roughly 1% of the compute required by previous zero‑shot baselines while preserving subject fidelity and motion quality.
By Daneul Kim, Jingxu Zhang, Wonjoon Jin, Sunghyun Cho, Qi Dai, Jaesik Park, Chong Luo
ViRDM is a new post‑training method for few‑step causal video generation that eliminates the need for a large teacher model and an online critic. By applying representation distribution matching (RDM) with a precomputed target distribution, a lightweight VAE decoder, and staged vector–Jacobian products, ViRDM overcomes memory, optimization, and temporal dynamics challenges. The approach reduces GPU memory usage and training time, achieving state‑of‑the‑art VBench performance with only 20 generator updates and 16 A100 GPU‑hours.
By Zichong Meng, Chongjian Ge, Chun-Hao P. Huang, Yang Zhou, Huaizu Jiang
arXiv:2607. 20628v1 Announce Type: cross Abstract: Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic training data, yet robust restoration is critical for downstream pipelines such as mobile imaging and 3D reconstruction.
By Renbiao Jin, Mingxin Yang, Yutian Chen, Junhao Zhuang, Xin Cai, Mulin Yu, Linning Xu, Wenxian Yu, Danping Zou, Shi Guo, Tianfan Xue
arXiv:2604.06655v2 Announce Type: replace
Abstract: Diffusion-based generative video compression offers a promising paradigm for low-bitrate reconstruction, but existing keyframe-based controllable a...
By Ding Ding, Daowen Li, Yixin Gao, Ruixiao Dong, Kai Li, Ying Chen, Li Li
LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.
By Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both...
arXiv:2606. 10450v1 Announce Type: cross Abstract: DiffC provides a principled way to reuse pre-trained diffusion models for lossy compression, but its encoding and decoding procedures remain slow because they require many discretized forward and reverse steps.
By Fuma Kimishima, Jinjia Zhou
TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.
By Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai