arXiv AI By Ziyang Ou

CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier

Read the original on arXiv AI →

The paper explores whether CLIP embeddings can detect AI-generated images by using a frozen CLIP model to extract visual embeddings and training lightweight classifiers on top. On the CIFAKE benchmark, the approach achieves 95% accuracy without language reasoning, and 85% accuracy after few-shot adaptation with 20% of the data. Certain image types, such as wide-angle photographs and oil paintings, remain challenging, highlighting unexplored difficulties in AI-generated image classification.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 27

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

MLLMCLIP introduces a heterogeneous distillation framework that transfers multimodal knowledge from a generative Multimodal Large Language Model (MLLM) teacher directly into a discriminative CLIP student, eliminating the need for synthetic hard negatives. The method uses an attention-based per-layer token selection and a CKA-based distillation loss to bridge architectural differences between the two models. As a result, MLLMCLIP achieves state‑of‑the‑art compositional accuracy and improves zero‑shot classification and image‑text retrieval performance.

By Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki, Junmo Kim, Yuki Mitsufuji
arXiv Computer Vision
2d ago

Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification

arXiv:2603.24528v2 Announce Type: replace Abstract: Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that...

By Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bart{\l}omiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer