OpenAI Blog

DALL·E: Creating images from text

We’ve trained a neural network called DALL·E that creates images from text captions for a wide range of concepts expressible in natural language.

OpenAI Blog
Jan 5, 2021

CLIP: Connecting text and images

We’re introducing a neural network called CLIP which efficiently learns visual concepts from natural language supervision. CLIP can be applied to any visual classification benchmark by simply providing the names of the visual categories to be recognized, similar to the “zero-shot” capabilities of GPT-2 and GPT-3.

arXiv Computer Vision
Aug 27

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

The paper introduces the CO-AID dataset, which captures systematic defects in state‑of‑the‑art text‑to‑image models when prompts involve complex composition such as multiple entities and attributes. Researchers manually curated 651 reference images across people, hand, object, and scene categories, edited ChatGPT‑generated prompts to emphasize compositional factors, and generated AI images with three T2I models. A subjective study with 29 participants produced multi‑label defect annotations, enabling training of a deep model that predicts defects and improves image generation.

By Ruoqi Hu, Chulin Zhao, Jiashuo Chang, Ramon Ruiz-Dolz, Hanhe Lin
arXiv Computer Vision
6d ago

Preserve-and-Compose Training for Composed Image Retrieval

The paper introduces Preserve-and-Compose Training (PACT) for composed image retrieval, a task where a query image is modified by a textual instruction while preserving visual content from a reference image. PACT learns from image–text–text triplets, using target captions for supervision and visual evidence from the source image to maintain relevant details, without requiring target images or gallery updates. The authors also propose Chord scoring, which blends target similarity with source-relative directional agreement in a frozen image space, and demonstrate that this combined approach yields strong retrieval performance across multiple zero-shot CIR benchmarks and various backbones.

By Sehyun Kwon
Hugging Face Trending Papers
Jun 17

Hierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning

Traditional image captioning methods often struggle to generate comprehensive, context-rich descriptions, especially for details not directly observable from visual cues. To overcome this, we propose a novel retrieval-augmented image captioning framework that generates captions with deeper insights, such as object attributes, event context, and underlying significance, by leveraging external knowledge.