OpenAI Blog

Multimodal neurons in artificial neural networks

We’ve discovered neurons in CLIP that respond to the same concept whether presented literally, symbolically, or conceptually. This may explain CLIP’s accuracy in classifying surprising visual renditions of concepts, and is also an important step toward understanding the associations and biases that CLIP and similar models learn.

arXiv Computer Vision
Aug 27

On the Separation of Human and AI-Generated Images in CLIP Embedding Space

The paper reports a new phenomenon in CLIP embeddings where human and AI‑generated paintings naturally separate along dominant principal directions without any supervised training. The authors investigate this separation by linking embedding directions back to image features using interpretable representations and gradient‑based inversion, finding that the separation is driven by distributed multiscale image structure rather than simple global or local statistics. They also show that small, imperceptible image perturbations can cause large displacements along these directions, highlighting a mismatch between CLIP’s visual evidence and human perception.

By Andrea Asperti
arXiv AI
Sep 2

Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures

The paper investigates whether neural networks exhibit conceptual separation, meaning that examples of the same concept cluster together and related concepts are closer in representation space. Using geometric and distributional analyses, the authors find that Convolutional Neural Networks (CNNs) produce coherent, semantically ordered representations for familiar ImageNet concepts, but this coherence weakens for unseen concepts and under domain shift. Large Language Models (LLMs) keep distinct domains well separated, bring related subdomains closer, yet lose distinction between ambiguous topics at both mean and covariance levels.

By Jaee Ponde, Roshni Agarwal, Subhashis Banerjee
OpenAI Blog
Jan 5, 2021

CLIP: Connecting text and images

We’re introducing a neural network called CLIP which efficiently learns visual concepts from natural language supervision. CLIP can be applied to any visual classification benchmark by simply providing the names of the visual categories to be recognized, similar to the “zero-shot” capabilities of GPT-2 and GPT-3.

Hugging Face Trending Papers
Jul 8

Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts

Claims about the universality of human concepts have been predominantly assessed through linguistic similarity across languages and cultures. However, words are effective as communication devices because they compress rich experiential variation into shared conventions, potentially obscuring hidden individual and cultural differences in how concepts are mentally represented.

arXiv AI
Sep 16

CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier

The paper explores whether CLIP embeddings can detect AI-generated images by using a frozen CLIP model to extract visual embeddings and training lightweight classifiers on top. On the CIFAKE benchmark, the approach achieves 95% accuracy without language reasoning, and 85% accuracy after few-shot adaptation with 20% of the data. Certain image types, such as wide-angle photographs and oil paintings, remain challenging, highlighting unexplored difficulties in AI-generated image classification.

By Ziyang Ou
arXiv AI
Aug 10

Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.

By Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy