Zero-shot image-to-text generation with BLIP-2
Related stories
Introducing ChatGPT Images 2.0
ChatGPT Images 2. 0 introduces a state-of-the-art image generation model with improved text rendering, multilingual support, and advanced visual reasoning.
Introducing 4o Image Generation
At OpenAI, we have long believed image generation should be a primary capability of our language models. That’s why we’ve built our most advanced image generator yet into GPT‑4o.
Open Preference Dataset for Text-to-Image Generation by the 🤗 Community
BLM-SGAN: Bidirectional Language Modeling for Semantic-Spatial Text-to-Image Generation
arXiv:2606. 08847v1 Announce Type: cross Abstract: Despite the success of image generation from text descriptions, it still faces challenges that are difficult to overcome in domains such as natural language processing (NLP) and computer vision (CV).
Training Design for Text-to-Image Models: Lessons from Ablations
Hierarchical text-conditional image generation with CLIP latents
Nexus: Structured Synergy for Efficient Text-to-Image Generation using Rectified Flow Model
Diffusion and flow matching models have made significant progress in text-to-image generation, yet high computation, quadratic complexity, and large memory footprint hinder high-resolution synthesis and edge deployment. We propose Nexus, which integrates sparse architecture, linear complexity, and low-bit quantization.
PRX Part 3 — Training a Text-to-Image Model in 24h!
Assisted Generation: a new direction toward low-latency text generation
Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech
arXiv:2606. 14750v1 Announce Type: cross Abstract: Recent advances in pixel-based text modeling show that representing text as images enables models to exploit visual cues for language understanding.
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
arXiv:2507. 18632v2 Announce Type: replace-cross Abstract: Zero-shot domain adaptation is a method for adapting a model to a target domain without utilizing target domain image data.