arXiv Computer Vision

PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers

arXiv Computer Vision
Sep 4

DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video Generation

The paper introduces DSAQuant, a quantization‑aware training framework tailored for video diffusion models (VDMs). It aligns quantization with the denoising stages of VDMs, using denoising‑stage‑oriented supervision during training and denoising‑stage gated guidance during inference to preserve structure while improving detail reconstruction. Experiments on Wan and CogVideoX models under aggressive W3A3 and W4A4 quantization settings show that DSAQuant outperforms state‑of‑the‑art QAT baselines, boosting VBench scores by up to 6.60 while maintaining strong text‑video alignment.

By Shuaiting Li, Zelin Gao, Haibin Shen, Yujun Shen, Haotong Qin, Yinghao Xu
arXiv Computer Vision
Sep 21

Quantization-Aware Kalman Estimation for Diffusion Sampling

The paper introduces QuAKE, a Quantization-Aware Kalman Estimator designed to correct errors in diffusion model sampling when using quantized denoisers. By treating sampling as an online estimation problem, QuAKE leverages the history of quantized outputs to recover full-precision estimates, updating a posterior in closed form at each step. The method is lightweight, plug‑and‑play, and works with any high‑order multistep ODE sampler, outperforming existing correction techniques on W4A4‑quantized text‑to‑image diffusion models.

By Qitan Shi, Cheng Jin, Jiawei Zhang, Yuantao Gu
arXiv Machine Learning
Aug 28

Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs

The paper investigates low‑bit quantization for Multimodal Large Language Models (MLLMs), showing that MXFP8 retains near‑lossless performance while 4‑bit formats like MXFP4 and HiF4 cause significant degradation. It identifies activation quantization as the main source of this loss and introduces Residual Fallback Quantization (RFQ), a lightweight framework that adds a quantized residual pathway to improve activation fidelity without architectural changes. Experiments on Wan2.2 and Qwen3‑VL demonstrate that RFQ recovers much of the performance gap to BF16 baselines across generation and reasoning tasks.

By Tanzila Rahman, Mehran Taghian Jazi, Yunke Peng, Zhuang Ma, Anandharaju Durai Raju, Yao Wang, Xing Huang, Hei Yi Mak, Shadan Golestan, Hoang Le, Yonghan Dong, Wei Guo, Yaoyuan Wang
arXiv Computer Vision
4d ago

FastVR: Efficient Streaming Video Restoration with One-Step Diffusion

FastVR is a streaming video restoration framework that uses a one‑step diffusion model to achieve strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. It addresses efficiency bottlenecks by combining a lightweight VAE with chunk‑wise causal attention, and improves inference speed and restoration quality through velocity consistency regularization and continuous trajectory learning during training. Experiments demonstrate that FastVR outperforms diffusion baselines in efficiency and achieves state‑of‑the‑art performance on both synthetic and real‑world benchmarks.

By Xiaoxu Chen, Qin Yang, Haoran Bai, Sibin Deng, Ying Chen