arXiv:2608.20515v1 Announce Type: new
Abstract: Generative video compression can recover rich visual details at low bitrates, but simultaneously achieving high temporal consistency and low inference...
By Wenzhuo Ma, Zhenzhong Chen
arXiv:2604.06655v2 Announce Type: replace
Abstract: Diffusion-based generative video compression offers a promising paradigm for low-bitrate reconstruction, but existing keyframe-based controllable a...
By Ding Ding, Daowen Li, Yixin Gao, Ruixiao Dong, Kai Li, Ying Chen, Li Li
Diffusion-based generative video compression has emerged as a promising paradigm to improve perceptual quality, where latent frames are required to be encoded efficiently while serving as denoising conditions. However, existing methods neither carefully design reference and quality structures during latent coding nor account for the impact of frame-level quality variation on denoising procedure, which limits coding efficiency and aggravates artifact propagation during generative reconstruction.
Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full se...
FastVR is a streaming video restoration framework that uses a one‑step diffusion model to achieve strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. It addresses efficiency bottlenecks by combining a lightweight VAE with chunk‑wise causal attention, and improves inference speed and restoration quality through velocity consistency regularization and continuous trajectory learning during training. Experiments demonstrate that FastVR outperforms diffusion baselines in efficiency and achieves state‑of‑the‑art performance on both synthetic and real‑world benchmarks.
By Xiaoxu Chen, Qin Yang, Haoran Bai, Sibin Deng, Ying Chen
arXiv:2512.23709v3 Announce Type: replace
Abstract: Diffusion-based video super-resolution (VSR) methods deliver strong perceptual quality but are often unsuitable for latency-sensitive scenarios due...
By Hau-Shiang Shiu, Chin-Yang Lin, Zhixiang Wang, Chi-Wei Hsiao, Po-Fan Yu, Yu-Chih Chen, Yu-Lun Liu
We present the Large Processing Model (LPM), a diffusion-based generative framework for photorealistic video restoration under complex, in-the-wild degradations. To our knowledge, LPM is the first generative video restoration model deployed at industrial scale.
arXiv:2607. 03803v1 Announce Type: cross Abstract: The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc.
By Xuyao Huang, Zelai Deng, Xu Wang, Xizhong Xiao, Zhijie Deng
The paper introduces VoRTeC, a video compression framework that leverages a foundational flow model to encode latent video representations compactly and predict their positions along flow trajectories. By integrating multi‑scale priors and avoiding access to flow‑matching network parameters, VoRTeC achieves one‑step decoding with high perceptual fidelity, while maintaining temporal consistency through tail‑frame reuse and prior caching. Experiments show a 58% reduction in bit consumption compared to prior diffusion‑based methods and a decoding speed increase ranging from 3 to 197 times, reaching 13 FPS at 720p and 32 FPS at 480p.
By Yichong Xia, Qinhong Wu, Qinhong Wu, Jinpeng Wang, Zeyuan Chen, Haoqian Wang
The paper introduces DSAQuant, a quantization‑aware training framework tailored for video diffusion models (VDMs). It aligns quantization with the denoising stages of VDMs, using denoising‑stage‑oriented supervision during training and denoising‑stage gated guidance during inference to preserve structure while improving detail reconstruction. Experiments on Wan and CogVideoX models under aggressive W3A3 and W4A4 quantization settings show that DSAQuant outperforms state‑of‑the‑art QAT baselines, boosting VBench scores by up to 6.60 while maintaining strong text‑video alignment.
By Shuaiting Li, Zelin Gao, Haibin Shen, Yujun Shen, Haotong Qin, Yinghao Xu
arXiv:2601. 11641v3 Announce Type: replace-cross Abstract: While Diffusion Transformers (DiTs) have achieved notable progress in video generation, this long-sequence generation task remains constrained by the quadratic complexity inherent to self-attention mechanisms, creating significant barriers to practical deployment.
By Yuxi Liu, Yipeng Hu, Zekun Zhang, Kunze Jiang, Kun Yuan
This article reviews recent diffusion‑based methods for generative lossy image compression, highlighting how these techniques encode a source into an embedding and use a diffusion model to iteratively refine the reconstruction during decoding. It discusses the role of auxiliary entropy models for transmitting the embedding, explores the use of diffusion models for information transmission via channel simulation, and frames the discussion within rate‑distortion‑perception theory, common randomness, and inverse‑problem connections. The review also identifies open challenges in the field.
By Yibo Yang, Stephan Mandt