Optimizing Stable Diffusion for Intel CPUs with NNCF and 🤗 Optimum
Related stories
Accelerating Stable Diffusion Inference on Intel CPUs
Using Stable Diffusion with Core ML on Apple Silicon
Swift 🧨Diffusers - Fast Stable Diffusion for Mac
Realizing Native INT8 Compute for Diffusion Transformers on Consumer GPUs: A Fused INT8 GEMM Kernel for Ideogram 4.0
arXiv:2606. 14598v1 Announce Type: new Abstract: Post-training INT8 (W8A8) quantization of diffusion transformers is widely deployed as a speed optimization, yet on consumer Ampere GPUs it is frequently slower than the FP8 and NF4 alternatives it is meant to beat.
Faster Stable Diffusion with Core ML on iPhone, iPad, and Mac
Instruction-tuning Stable Diffusion with InstructPix2Pix
Performance Analysis and Optimization of 3D Generative Diffusion Models across GPU Architectures
arXiv:2606. 19365v1 Announce Type: new Abstract: Diffusion models have become essential for high-fidelity 3D MRI synthesis, yet their deployment remains constrained by substantial GPU resource demands arising from hundreds of U-Net evaluations per sample and a highly heterogeneous kernel behavior.
Accelerating PyTorch distributed fine-tuning with Intel technologies
Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational Hardening
We report operational experience full-fine-tuning a 32. 76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) -- among the first published field accounts on this accelerator.
Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational Hardening
arXiv:2608. 05944v1 Announce Type: cross Abstract: We report operational experience full-fine-tuning a 32.