Training Design for Text-to-Image Models: Lessons from Ablations
Related stories
A Dive into Text-to-Video Models
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
arXiv:2505. 16915v3 Announce Type: replace-cross Abstract: While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the long, detailed prompts required for professional applications.
Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation
arXiv:2608. 14172v1 Announce Type: cross Abstract: Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
arXiv:2603. 00133v2 Announce Type: replace-cross Abstract: Generative models have been shown to "memorize" certain training data, leading to verbatim or near-verbatim generating images, which may cause privacy concerns or copyright infringement.
Welcome aMUSEd: Efficient Text-to-Image Generation
Zero-shot image-to-text generation with BLIP-2
Introducing ConTextual: How well can your Multimodal model jointly reason over text and image in text-rich scenes?
Open Preference Dataset for Text-to-Image Generation by the 🤗 Community
Abra: Scaling Diffusion Image Training
arXiv:2608. 17286v1 Announce Type: new Abstract: Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation.
DALL·E: Creating images from text
We’ve trained a neural network called DALL·E that creates images from text captions for a wide range of concepts expressible in natural language.