arXiv:2609.33384v2 Announce Type: replace
Abstract: Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction err...
By Yutong Wang, Xingtong Ge, Enhuai Liu, Yunke Wang, Tianfan Xue, Xinyuan Chen, Chang Xu
arXiv:2607. 21446v1 Announce Type: new Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent.
By Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann
arXiv:2607. 03803v1 Announce Type: cross Abstract: The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc.
By Xuyao Huang, Zelai Deng, Xu Wang, Xizhong Xiao, Zhijie Deng
The paper introduces DSAQuant, a quantization‑aware training framework tailored for video diffusion models (VDMs). It aligns quantization with the denoising stages of VDMs, using denoising‑stage‑oriented supervision during training and denoising‑stage gated guidance during inference to preserve structure while improving detail reconstruction. Experiments on Wan and CogVideoX models under aggressive W3A3 and W4A4 quantization settings show that DSAQuant outperforms state‑of‑the‑art QAT baselines, boosting VBench scores by up to 6.60 while maintaining strong text‑video alignment.
By Shuaiting Li, Zelin Gao, Haibin Shen, Yujun Shen, Haotong Qin, Yinghao Xu
arXiv:2606. 00094v1 Announce Type: cross Abstract: Image generative models aim to sample data points from the underlying data manifold, a task that requires learning and decoding a dense, low-dimensional, and compact parameterization space.
By Duoduo Xue, Zhiyu Zhu, Junhui Hou
arXiv:2502. 10389v2 Announce Type: replace-cross Abstract: Diffusion models (DMs) have become the leading choice for generative tasks across diverse domains.
By Ziming Liu, Yifan Yang, Chengruidong Zhang, Yiqi Zhang, Lili Qiu, Yang You, Yuqing Yang
arXiv:2602. 13357v3 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art performance in high-fidelity image and video generation but suffer from expensive inference due to their iterative denoising structure.
By Dong Liu, Yanxuan Yu, Ben Lengerich, Ying Nian Wu
arXiv:2603. 14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility?
By Chujun Tang, Lei Zhong, Fangqiang Ding
arXiv:2610.00930v1 Announce Type: new
Abstract: Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating...
By Mingrun Jiang, Yuejia Liu, Zishan Shao, Ting Jiang, Qinsi Wang, Hancheng Ye, Yixiao Wang, Rui-Feng Wang, Kangning Cui, Yixuan Chen, Fan Yang, Xiang Cheng, Hai Li, Yiran Chen
arXiv:2604.06161v3 Announce Type: replace-cross
Abstract: Most digital videos are stored in 8-bit low dynamic range (LDR) formats, where much of the original high dynamic range (HDR) scene radiance i...
By Zhengming Yu, Li Ma, Mingming He, Leo Isikdogan, Yuancheng Xu, Dmitriy Smirnov, Pablo Salamanca, Dao Mi, Pablo Delgado, Ning Yu, Julien Philip, Xin Li, Wenping Wang, Paul Debevec
arXiv:2608. 11045v1 Announce Type: new Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals.
By He-Yen Hsieh, H. T. Kung
Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different degradations using a single model. Most recent methods adapt large pretrained text-to-image (T2I) latent diffusion models for their strong capacity and generative priors.