Welcome aMUSEd: Efficient Text-to-Image Generation
Related stories
Hierarchical text-conditional image generation with CLIP latents
Open Preference Dataset for Text-to-Image Generation by the 🤗 Community
Assisted Generation: a new direction toward low-latency text generation
Introducing 4o Image Generation
At OpenAI, we have long believed image generation should be a primary capability of our language models. That’s why we’ve built our most advanced image generator yet into GPT‑4o.
Introducing ChatGPT Images 2.0
ChatGPT Images 2. 0 introduces a state-of-the-art image generation model with improved text rendering, multilingual support, and advanced visual reasoning.
Evaluating Reasoning Fidelity in Visual Text Generation
arXiv:2606. 04479v1 Announce Type: cross Abstract: Recent text-to-image (T2I) models can render highly legible and well-structured text within images, enabling applications including document generation and slide generation.
PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation
Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where models must align outputs with a user's implicit visual preferences based on a few historically preferred images and a short prompt.
Nexus: Structured Synergy for Efficient Text-to-Image Generation using Rectified Flow Model
Diffusion and flow matching models have made significant progress in text-to-image generation, yet high computation, quadratic complexity, and large memory footprint hinder high-resolution synthesis and edge deployment. We propose Nexus, which integrates sparse architecture, linear complexity, and low-bit quantization.
Emotion-Aware Image Generation from Korean Diary Text via LLM-based Prompt Translation and LoRA Fine-Tuning
arXiv:2606. 05816v1 Announce Type: cross Abstract: T2I models cannot effectively capture sentiment from various types of text, including diaries, as they primarily focus on visual object-related patterns rather than contextual emotional understanding.
Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation
arXiv:2603. 12506v2 Announce Type: replace-cross Abstract: Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise.
BLM-SGAN: Bidirectional Language Modeling for Semantic-Spatial Text-to-Image Generation
arXiv:2606. 08847v1 Announce Type: cross Abstract: Despite the success of image generation from text descriptions, it still faces challenges that are difficult to overcome in domains such as natural language processing (NLP) and computer vision (CV).