Optimizing Stable Diffusion for Intel CPUs with NNCF and đ¤ Optimum
Related stories
Accelerating Stable Diffusion Inference on Intel CPUs
Using Stable Diffusion with Core ML on Apple Silicon
Swift đ§¨Diffusers - Fast Stable Diffusion for Mac
Realizing Native INT8 Compute for Diffusion Transformers on Consumer GPUs: A Fused INT8 GEMM Kernel for Ideogram 4.0
arXiv:2606. 14598v1 Announce Type: new Abstract: Post-training INT8 (W8A8) quantization of diffusion transformers is widely deployed as a speed optimization, yet on consumer Ampere GPUs it is frequently slower than the FP8 and NF4 alternatives it is meant to beat.
Faster Stable Diffusion with Core ML on iPhone, iPad, and Mac
Instruction-tuning Stable Diffusion with InstructPix2Pix
Performance Analysis and Optimization of 3D Generative Diffusion Models across GPU Architectures
arXiv:2606. 19365v1 Announce Type: new Abstract: Diffusion models have become essential for high-fidelity 3D MRI synthesis, yet their deployment remains constrained by substantial GPU resource demands arising from hundreds of U-Net evaluations per sample and a highly heterogeneous kernel behavior.
Format-Aware Fusion for Fast FP4 Pretraining
The paper introduces formatâaware fusion, a method that coâdesigns quantization producers with their scale domains and consumer layouts to fully exploit fourâbit floatingâpoint (FP4) Tensor Cores. Using this approach, the authors pretrain the Llamaâ3âfamily 8B model on 160âŻbillion tokens, achieving up to 37.9âŻK tokens/s/GPUâsignificantly higher than standard bfloat16 or Transformer Engine FP4 baselines. The study demonstrates that FP4 performance depends on the interplay of scaling, operand packing, layout, and execution path, with downstream task rankings diverging from trainingâloss rankings.
Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware
Masked diffusion language models (dLLMs) promise faster text generation by denoising multiple tokens simultaneously, yet their realâworld serving behavior has been largely unexamined. Using LLaDAâ8BâInstruct on a single NVIDIA H200 GPU, the study finds that request difficulty is discretized into 11 stepâcount levels, shortâbudget benchmarks underestimate serving variance, and only 24% of singleârequest time is GPU computation, with batching mainly reducing CPU dispatch overhead. The authors also demonstrate that output quality remains stable across batch sizes and propose a batchâtimeout rule for synchronized batching under Poisson arrivals.
Accelerating PyTorch distributed fine-tuning with Intel technologies
Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning
The study investigates how the rank of LowâRank Adaptation (LoRA) affects diffusion model fineâtuning on CIFARâ10 using a DDPM UâNet. Experiments with ranks 2, 4, 8, 16, and 32 show that moderate ranksâparticularly rankâŻ4âyield the best FID scores while keeping trainable parameters, runtime, and GPU memory low. Higher ranks offer only marginal improvements at a higher computational cost, suggesting that smallâtoâmoderate ranks are efficient defaults for fixed training budgets.