Hierarchical text-conditional image generation with CLIP latents
Related stories
Open Preference Dataset for Text-to-Image Generation by the 🤗 Community
Introducing Würstchen: Fast Diffusion for Image Generation
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
arXiv:2607. 19344v1 Announce Type: cross Abstract: Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone.
Zero-shot image-to-text generation with BLIP-2
Emotion-Aware Image Generation from Korean Diary Text via LLM-based Prompt Translation and LoRA Fine-Tuning
arXiv:2606. 05816v1 Announce Type: cross Abstract: T2I models cannot effectively capture sentiment from various types of text, including diaries, as they primarily focus on visual object-related patterns rather than contextual emotional understanding.
Better Source, Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching
arXiv:2602. 05951v2 Announce Type: replace-cross Abstract: Flow matching has recently emerged as a promising alternative to diffusion-based generative models, particularly for text-to-image generation.
Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation
arXiv:2603. 12506v2 Announce Type: replace-cross Abstract: Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise.
PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation
Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where models must align outputs with a user's implicit visual preferences based on a few historically preferred images and a short prompt.
Introducing ChatGPT Images 2.0
ChatGPT Images 2. 0 introduces a state-of-the-art image generation model with improved text rendering, multilingual support, and advanced visual reasoning.