arXiv AI By Donghyun Lee, Jitesh Chavan, Duy Nguyen, Sam Huang, Liming Jiang, Priyadarshini Panda, Timo Mertens, Saurabh Shukla

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

Read the original on arXiv AI →

arXiv:2607. 02461v1 Announce Type: cross Abstract: Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make inference expensive.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 24

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

arXiv:2607. 21446v1 Announce Type: new Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent.

By Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann
arXiv Computer Vision
Sep 4

DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video Generation

The paper introduces DSAQuant, a quantization‑aware training framework tailored for video diffusion models (VDMs). It aligns quantization with the denoising stages of VDMs, using denoising‑stage‑oriented supervision during training and denoising‑stage gated guidance during inference to preserve structure while improving detail reconstruction. Experiments on Wan and CogVideoX models under aggressive W3A3 and W4A4 quantization settings show that DSAQuant outperforms state‑of‑the‑art QAT baselines, boosting VBench scores by up to 6.60 while maintaining strong text‑video alignment.

By Shuaiting Li, Zelin Gao, Haibin Shen, Yujun Shen, Haotong Qin, Yinghao Xu