Fine-tuning Stable Diffusion models on Intel CPUs
Related stories
Optimizing Stable Diffusion for Intel CPUs with NNCF and 🤗 Optimum
Using Stable Diffusion with Core ML on Apple Silicon
Instruction-tuning Stable Diffusion with InstructPix2Pix
Faster Stable Diffusion with Core ML on iPhone, iPad, and Mac
Using LoRA for Efficient Stable Diffusion Fine-Tuning
Swift 🧨Diffusers - Fast Stable Diffusion for Mac
Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning
The study investigates how the rank of Low‑Rank Adaptation (LoRA) affects diffusion model fine‑tuning on CIFAR‑10 using a DDPM U‑Net. Experiments with ranks 2, 4, 8, 16, and 32 show that moderate ranks—particularly rank 4—yield the best FID scores while keeping trainable parameters, runtime, and GPU memory low. Higher ranks offer only marginal improvements at a higher computational cost, suggesting that small‑to‑moderate ranks are efficient defaults for fixed training budgets.
Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware
Masked diffusion language models (dLLMs) promise faster text generation by denoising multiple tokens simultaneously, yet their real‑world serving behavior has been largely unexamined. Using LLaDA‑8B‑Instruct on a single NVIDIA H200 GPU, the study finds that request difficulty is discretized into 11 step‑count levels, short‑budget benchmarks underestimate serving variance, and only 24% of single‑request time is GPU computation, with batching mainly reducing CPU dispatch overhead. The authors also demonstrate that output quality remains stable across batch sizes and propose a batch‑timeout rule for synchronized batching under Poisson arrivals.
Performance Analysis and Optimization of 3D Generative Diffusion Models across GPU Architectures
arXiv:2606. 19365v1 Announce Type: new Abstract: Diffusion models have become essential for high-fidelity 3D MRI synthesis, yet their deployment remains constrained by substantial GPU resource demands arising from hundreds of U-Net evaluations per sample and a highly heterogeneous kernel behavior.
ASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory Hardware
arXiv:2609.00955v1 Announce Type: new Abstract: Diffusion models achieve strong image generation quality but incur high iterative denoising costs. Analog compute-in-memory (CIM) can accelerate matrix...
Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation Quality
arXiv:2608. 02575v1 Announce Type: new Abstract: Diffusion models rely on stochastic inputs, yet on finite-precision hardware, the "randomness" they consume is realized as deterministic numerical orbits generated by pseudorandom rules.