arXiv Machine Learning By Ali Asaria, Tony Salomone, Deep Gandhi

Realizing Native INT8 Compute for Diffusion Transformers on Consumer GPUs: A Fused INT8 GEMM Kernel for Ideogram 4.0

Read the original on arXiv Machine Learning →

arXiv:2606. 14598v1 Announce Type: new Abstract: Post-training INT8 (W8A8) quantization of diffusion transformers is widely deployed as a speed optimization, yet on consumer Ampere GPUs it is frequently slower than the FP8 and NF4 alternatives it is meant to beat.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.