Instruction-tuning Stable Diffusion with InstructPix2Pix
Related stories
Accelerating Stable Diffusion Inference on Intel CPUs
Optimizing Stable Diffusion for Intel CPUs with NNCF and 🤗 Optimum
Using Stable Diffusion with Core ML on Apple Silicon
Using LoRA for Efficient Stable Diffusion Fine-Tuning
Make LLM Fine-tuning 2x faster with Unsloth and 🤗 TRL
Beyond Attention Masks: Instruction Anchoring for Efficient In-Context Diffusion Generation
The paper introduces AnchorCache, a parameter‑free token‑layout and attention‑mask design that decouples reference tokens from the target in in‑context diffusion transformers. By inserting static text anchors, the method conditions reference representations on the instruction during cache construction, enabling exact key‑value reuse across denoising steps. To restore quality lost by this structural change, the authors employ teacher‑forced velocity distillation followed by a brief on‑policy stage, achieving full‑attention quality while delivering up to 6.40× speedup in diffusion transformer inference across image, speech, and video benchmarks.
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
ASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory Hardware
arXiv:2609.00955v1 Announce Type: new Abstract: Diffusion models achieve strong image generation quality but incur high iterative denoising costs. Analog compute-in-memory (CIM) can accelerate matrix...
Performance Analysis and Optimization of 3D Generative Diffusion Models across GPU Architectures
arXiv:2606. 19365v1 Announce Type: new Abstract: Diffusion models have become essential for high-fidelity 3D MRI synthesis, yet their deployment remains constrained by substantial GPU resource demands arising from hundreds of U-Net evaluations per sample and a highly heterogeneous kernel behavior.
RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention
arXiv:2608. 08081v1 Announce Type: cross Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand.
Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights
arXiv:2610.02598v1 Announce Type: new Abstract: Deploying LLMs on consumer-grade GPUs with insufficient memory to hold their weights can result in prohibitively slow inference, because decoding repea...