Stable Diffusion XL on Mac with Advanced Core ML Quantization
Related stories
Swift 🧨Diffusers - Fast Stable Diffusion for Mac
Using Stable Diffusion with Core ML on Apple Silicon
🧨 Accelerating Stable Diffusion XL Inference with JAX on Cloud TPU v5e
Accelerating Stable Diffusion Inference on Intel CPUs
Bringing Nunchaku 4-bit Diffusion Inference to Diffusers
Quanto: a PyTorch quantization backend for Optimum
Quantization-Aware Kalman Estimation for Diffusion Sampling
The paper introduces QuAKE, a Quantization-Aware Kalman Estimator designed to correct errors in diffusion model sampling when using quantized denoisers. By treating sampling as an online estimation problem, QuAKE leverages the history of quantized outputs to recover full-precision estimates, updating a posterior in closed form at each step. The method is lightweight, plug‑and‑play, and works with any high‑order multistep ODE sampler, outperforming existing correction techniques on W4A4‑quantized text‑to‑image diffusion models.
Exploring Quantization Backends in Diffusers
Memory-efficient Diffusion Transformers with Quanto and Diffusers
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
arXiv:2601. 22813v2 Announce Type: replace Abstract: The NVFP4 lower-precision format, supported in hardware by NVIDIA Blackwell GPUs, promises to allow, for the first time, end-to-end fully-quantized pre-training of massive models such as LLMs.
KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers
arXiv:2607. 21446v1 Announce Type: new Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent.