Hugging Face Blog

Training Design for Text-to-Image Models: Lessons from Ablations

arXiv Computer Vision
Sep 7

How far can we go with ImageNet for Text-to-Image generation?

The paper argues that large-scale text‑to‑image models can be trained effectively on ImageNet, provided the dataset is enriched with carefully crafted text and image augmentations. Using this approach, the authors match the performance of state‑of‑the‑art models such as FLUX, surpassing SD3 on GenEval by +5 points and SDXL on DPGBench by +12, while employing only 1/1000th of the training images and significantly fewer parameters. The method requires just 500 hours of H100 GPU time, making it a more reproducible and accessible alternative to massive web‑scraped datasets.

By L. Degeorge, A. Ghosh, N. Dufour, D. Picard, V. Kalogeiton