Training Design for Text-to-Image Models: Lessons from Ablations
Related stories
A Dive into Text-to-Video Models
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
arXiv:2505. 16915v3 Announce Type: replace-cross Abstract: While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the long, detailed prompts required for professional applications.
Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation
arXiv:2608. 14172v1 Announce Type: cross Abstract: Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
arXiv:2603. 00133v2 Announce Type: replace-cross Abstract: Generative models have been shown to "memorize" certain training data, leading to verbatim or near-verbatim generating images, which may cause privacy concerns or copyright infringement.
Anatomy-Grounded Weakly Supervised Prompt Tuning for Chest X-ray Latent Diffusion Models
arXiv:2506.10633v2 Announce Type: replace Abstract: Latent Diffusion Models have shown remarkable results in text-guided image synthesis in recent years. In the domain of natural (RGB) images, recent...
Welcome aMUSEd: Efficient Text-to-Image Generation
How far can we go with ImageNet for Text-to-Image generation?
The paper argues that large-scale text‑to‑image models can be trained effectively on ImageNet, provided the dataset is enriched with carefully crafted text and image augmentations. Using this approach, the authors match the performance of state‑of‑the‑art models such as FLUX, surpassing SD3 on GenEval by +5 points and SDXL on DPGBench by +12, while employing only 1/1000th of the training images and significantly fewer parameters. The method requires just 500 hours of H100 GPU time, making it a more reproducible and accessible alternative to massive web‑scraped datasets.
Zero-shot image-to-text generation with BLIP-2
POET: Preference Optimization for Enhanced Text-to-Image Generation
arXiv:2510.12041v3 Announce Type: replace Abstract: Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecifie...