PRX Part 3 — Training a Text-to-Image Model in 24h!
Related stories
A Dive into Text-to-Video Models
How far can we go with ImageNet for Text-to-Image generation?
The paper argues that large-scale text‑to‑image models can be trained effectively on ImageNet, provided the dataset is enriched with carefully crafted text and image augmentations. Using this approach, the authors match the performance of state‑of‑the‑art models such as FLUX, surpassing SD3 on GenEval by +5 points and SDXL on DPGBench by +12, while employing only 1/1000th of the training images and significantly fewer parameters. The method requires just 500 hours of H100 GPU time, making it a more reproducible and accessible alternative to massive web‑scraped datasets.
Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models
The paper introduces an adaptive step schedule controller for text‑to‑image diffusion models, allowing the number of denoising steps to vary based on the complexity of the input prompt. By mixing step schedules of different sizes and monitoring error discrepancies at each timestep, the method switches schedules to maintain image quality while reducing inference time. Experiments on COCO and DiffusionDB demonstrate that this approach achieves faster generation without sacrificing visual fidelity.
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
arXiv:2505. 16915v3 Announce Type: replace-cross Abstract: While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the long, detailed prompts required for professional applications.
Welcome aMUSEd: Efficient Text-to-Image Generation
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
arXiv:2603. 00133v2 Announce Type: replace-cross Abstract: Generative models have been shown to "memorize" certain training data, leading to verbatim or near-verbatim generating images, which may cause privacy concerns or copyright infringement.
Unified Text-Image Generation with Weakness-Targeted Post-Training
The paper introduces a post‑training approach that enables a single inference process to transition from text reasoning to image synthesis, eliminating the need for explicit modality switching. Using the 14B BAGEL model, the authors demonstrate that targeted post‑training data and reward‑weighted training improve multimodal image generation across four independent T2I benchmarks. The study highlights the benefits of joint text‑image generation and strategic data selection for enhancing T2I performance.
BLM-SGAN: Bidirectional Language Modeling for Semantic-Spatial Text-to-Image Generation
arXiv:2606. 08847v1 Announce Type: cross Abstract: Despite the success of image generation from text descriptions, it still faces challenges that are difficult to overcome in domains such as natural language processing (NLP) and computer vision (CV).
Adversarial Learning of Classifier-Free Guidance Schedules
arXiv:2608. 14038v1 Announce Type: new Abstract: Modern text-to-image diffusion models rely on classifier-free guidance (CFG) to achieve high image fidelity and text alignment.
Introducing 4o Image Generation
At OpenAI, we have long believed image generation should be a primary capability of our language models. That’s why we’ve built our most advanced image generator yet into GPT‑4o.
AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow
AcFlow introduces an inference‑time controller for text‑to‑image diffusion transformers that transports intermediate layer activations through a learned, concept‑conditioned velocity field while keeping the base model frozen. The method allows fine‑grained style intensity control and suppression of unwanted concepts, achieving superior style–content trade‑offs compared to baselines and generalizing to unseen concepts without per‑concept fitting. Experiments demonstrate improved style alignment and qualitative suppression of diverse concepts where direct prompting fails.