arXiv Computer Vision
Sep 7

How far can we go with ImageNet for Text-to-Image generation?

The paper argues that large-scale text‑to‑image models can be trained effectively on ImageNet, provided the dataset is enriched with carefully crafted text and image augmentations. Using this approach, the authors match the performance of state‑of‑the‑art models such as FLUX, surpassing SD3 on GenEval by +5 points and SDXL on DPGBench by +12, while employing only 1/1000th of the training images and significantly fewer parameters. The method requires just 500 hours of H100 GPU time, making it a more reproducible and accessible alternative to massive web‑scraped datasets.

By L. Degeorge, A. Ghosh, N. Dufour, D. Picard, V. Kalogeiton
arXiv AI
Sep 16

Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models

The paper introduces an adaptive step schedule controller for text‑to‑image diffusion models, allowing the number of denoising steps to vary based on the complexity of the input prompt. By mixing step schedules of different sizes and monitoring error discrepancies at each timestep, the method switches schedules to maintain image quality while reducing inference time. Experiments on COCO and DiffusionDB demonstrate that this approach achieves faster generation without sacrificing visual fidelity.

By Kuluhan Binici, Cihan Acar, Shivam Aggarwal, Siying Liu, Tulika Mitra