Introducing Würstchen: Fast Diffusion for Image Generation
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
arXiv:2606. 04299v1 Announce Type: cross Abstract: We consider the problem of generating images whose internal structure -- defined by the distribution of patches across multiple scales -- matches that of a single reference image.
The paper introduces an adaptive step schedule controller for text‑to‑image diffusion models, allowing the number of denoising steps to vary based on the complexity of the input prompt. By mixing step schedules of different sizes and monitoring error discrepancies at each timestep, the method switches schedules to maintain image quality while reducing inference time. Experiments on COCO and DiffusionDB demonstrate that this approach achieves faster generation without sacrificing visual fidelity.
arXiv:2609.37654v1 Announce Type: new Abstract: We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-...
Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes slow generation.
This article reviews recent diffusion‑based methods for generative lossy image compression, highlighting how these techniques encode a source into an embedding and use a diffusion model to iteratively refine the reconstruction during decoding. It discusses the role of auxiliary entropy models for transmitting the embedding, explores the use of diffusion models for information transmission via channel simulation, and frames the discussion within rate‑distortion‑perception theory, common randomness, and inverse‑problem connections. The review also identifies open challenges in the field.