arXiv AI
Sep 16

Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models

The paper introduces an adaptive step schedule controller for text‑to‑image diffusion models, allowing the number of denoising steps to vary based on the complexity of the input prompt. By mixing step schedules of different sizes and monitoring error discrepancies at each timestep, the method switches schedules to maintain image quality while reducing inference time. Experiments on COCO and DiffusionDB demonstrate that this approach achieves faster generation without sacrificing visual fidelity.

By Kuluhan Binici, Cihan Acar, Shivam Aggarwal, Siying Liu, Tulika Mitra
OpenAI Blog
Jun 20, 2024

Consistency Models

Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes slow generation.

arXiv Machine Learning
Sep 24

Advances in Diffusion-Based Generative Compression

This article reviews recent diffusion‑based methods for generative lossy image compression, highlighting how these techniques encode a source into an embedding and use a diffusion model to iteratively refine the reconstruction during decoding. It discusses the role of auxiliary entropy models for transmitting the embedding, explores the use of diffusion models for information transmission via channel simulation, and frames the discussion within rate‑distortion‑perception theory, common randomness, and inverse‑problem connections. The review also identifies open challenges in the field.

By Yibo Yang, Stephan Mandt