OpenAI Blog

CLIP: Connecting text and images

Read the original on OpenAI Blog →

We’re introducing a neural network called CLIP which efficiently learns visual concepts from natural language supervision. CLIP can be applied to any visual classification benchmark by simply providing the names of the visual categories to be recognized, similar to the “zero-shot” capabilities of GPT-2 and GPT-3.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

Hugging Face Trending Papers
Jul 15

Fine-grained CLIP fine-tuning with self-annotated region alignment

Contrastive Language-Image Pre-training (CLIP) has been shown to have limitations in its fine-grained dense feature representation, due to its pre-training focusing on matching the whole image to a text description. Considering the large data and computational burden in pre-training a vision-language model from scratch, a series of works aim to enhance the fine-grained ability of CLIP through a fine-tuning scheme.

arXiv AI
Sep 16

CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier

The paper explores whether CLIP embeddings can detect AI-generated images by using a frozen CLIP model to extract visual embeddings and training lightweight classifiers on top. On the CIFAKE benchmark, the approach achieves 95% accuracy without language reasoning, and 85% accuracy after few-shot adaptation with 20% of the data. Certain image types, such as wide-angle photographs and oil paintings, remain challenging, highlighting unexplored difficulties in AI-generated image classification.

By Ziyang Ou
arXiv Computer Vision
6d ago

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

ProCAP introduces a probabilistic cross-attentive prompt learning framework for vision-language models like CLIP, enabling improved cross-modal interaction without updating the backbone. It jointly learns visual and textual prompt tokens, linking them via stacked bidirectional multi-head cross-attention to refine each branch across prompt depth. The method incorporates Gaussian parameterization of prompt tokens, lightweight KL and L2 regularization, and a compact symmetric InfoNCE head to align image features with class-level text representations, achieving strong few-shot base-to-novel performance and competitive transfer results across multiple datasets and benchmarks.

By Hiwa Azeez Abbas, Fatemeh Daneshfar, Moloud Abdar
arXiv Computer Vision
6d ago

Preserve-and-Compose Training for Composed Image Retrieval

The paper introduces Preserve-and-Compose Training (PACT) for composed image retrieval, a task where a query image is modified by a textual instruction while preserving visual content from a reference image. PACT learns from image–text–text triplets, using target captions for supervision and visual evidence from the source image to maintain relevant details, without requiring target images or gallery updates. The authors also propose Chord scoring, which blends target similarity with source-relative directional agreement in a frozen image space, and demonstrate that this combined approach yields strong retrieval performance across multiple zero-shot CIR benchmarks and various backbones.

By Sehyun Kwon
arXiv Computer Vision
3d ago

Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification

arXiv:2603.24528v2 Announce Type: replace Abstract: Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that...

By Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bart{\l}omiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer