Open Preference Dataset for Text-to-Image Generation by the đ¤ Community
Read the original on Hugging Face Blog âThe Flow has not summarised this story yet â read it at Hugging Face Blog.
The Flow has not summarised this story yet â read it at Hugging Face Blog.
Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where models must align outputs with a user's implicit visual preferences based on a few historically preferred images and a short prompt.
At OpenAI, we have long believed image generation should be a primary capability of our language models. Thatâs why weâve built our most advanced image generator yet into GPTâ4o.
The paper argues that large-scale textâtoâimage models can be trained effectively on ImageNet, provided the dataset is enriched with carefully crafted text and image augmentations. Using this approach, the authors match the performance of stateâofâtheâart models such as FLUX, surpassing SD3 on GenEval by +5 points and SDXL on DPGBench by +12, while employing only 1/1000th of the training images and significantly fewer parameters. The method requires just 500 hours of H100 GPU time, making it a more reproducible and accessible alternative to massive webâscraped datasets.
arXiv:2606. 05816v1 Announce Type: cross Abstract: T2I models cannot effectively capture sentiment from various types of text, including diaries, as they primarily focus on visual object-related patterns rather than contextual emotional understanding.
arXiv:2510.12041v3 Announce Type: replace Abstract: Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecifie...