arXiv:2607. 02461v1 Announce Type: cross Abstract: Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make inference expensive.
By Donghyun Lee, Jitesh Chavan, Duy Nguyen, Sam Huang, Liming Jiang, Priyadarshini Panda, Timo Mertens, Saurabh Shukla
The paper introduces DSAQuant, a quantization‑aware training framework tailored for video diffusion models (VDMs). It aligns quantization with the denoising stages of VDMs, using denoising‑stage‑oriented supervision during training and denoising‑stage gated guidance during inference to preserve structure while improving detail reconstruction. Experiments on Wan and CogVideoX models under aggressive W3A3 and W4A4 quantization settings show that DSAQuant outperforms state‑of‑the‑art QAT baselines, boosting VBench scores by up to 6.60 while maintaining strong text‑video alignment.
By Shuaiting Li, Zelin Gao, Haibin Shen, Yujun Shen, Haotong Qin, Yinghao Xu
arXiv:2608. 11045v1 Announce Type: new Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals.
By He-Yen Hsieh, H. T. Kung
The paper introduces QuAKE, a Quantization-Aware Kalman Estimator designed to correct errors in diffusion model sampling when using quantized denoisers. By treating sampling as an online estimation problem, QuAKE leverages the history of quantized outputs to recover full-precision estimates, updating a posterior in closed form at each step. The method is lightweight, plug‑and‑play, and works with any high‑order multistep ODE sampler, outperforming existing correction techniques on W4A4‑quantized text‑to‑image diffusion models.
By Qitan Shi, Cheng Jin, Jiawei Zhang, Yuantao Gu
Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full se...
arXiv:2608. 13932v1 Announce Type: new Abstract: Iterative Generative Models (IGMs) span autoregressive and diffusion paradigms, and hybrid variants that couple them can achieve remarkable image-generation fidelity.
By Jing Gao, Junyi Wu, Wei Wang, Yan Yan, Yao Zhao
arXiv:2606. 00658v1 Announce Type: cross Abstract: Large video diffusion models achieve strong visual quality but remain expensive to deploy because each sample requires many denoising steps and a large resident parameter footprint.
By Jinyang Du, Shenghao Jin, Ziqian Xu, Ruihao Gong, Shiqiao Gu, Yang Yong, Jinyang Guo, Xianglong Liu
The paper investigates low‑bit quantization for Multimodal Large Language Models (MLLMs), showing that MXFP8 retains near‑lossless performance while 4‑bit formats like MXFP4 and HiF4 cause significant degradation. It identifies activation quantization as the main source of this loss and introduces Residual Fallback Quantization (RFQ), a lightweight framework that adds a quantized residual pathway to improve activation fidelity without architectural changes. Experiments on Wan2.2 and Qwen3‑VL demonstrate that RFQ recovers much of the performance gap to BF16 baselines across generation and reasoning tasks.
By Tanzila Rahman, Mehran Taghian Jazi, Yunke Peng, Zhuang Ma, Anandharaju Durai Raju, Yao Wang, Xing Huang, Hei Yi Mak, Shadan Golestan, Hoang Le, Yonghan Dong, Wei Guo, Yaoyuan Wang
FastVR is a streaming video restoration framework that uses a one‑step diffusion model to achieve strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. It addresses efficiency bottlenecks by combining a lightweight VAE with chunk‑wise causal attention, and improves inference speed and restoration quality through velocity consistency regularization and continuous trajectory learning during training. Experiments demonstrate that FastVR outperforms diffusion baselines in efficiency and achieves state‑of‑the‑art performance on both synthetic and real‑world benchmarks.
By Xiaoxu Chen, Qin Yang, Haoran Bai, Sibin Deng, Ying Chen
arXiv:2603.17825v2 Announce Type: replace
Abstract: In this work, we study the role of Massive Activations (MAs), which are rare, high-magnitude spikes confined to a few fixed hidden dimensions in vi...
By Xianhang Cheng, Yujian Zheng, Zhenyu Xie, Tingting Liao, Hao Li
arXiv:2609.30820v1 Announce Type: cross
Abstract: Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes...
By Nux Li
arXiv:2603. 14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility?
By Chujun Tang, Lei Zhong, Fangqiang Ding