arXiv:2606. 19259v2 Announce Type: replace-cross Abstract: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information.
By Yijin Wang, Shuyi Wang, Wenhan Zhang, Yuqi Ouyang
arXiv:2606. 19259v1 Announce Type: cross Abstract: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information.
By Yijin Wang, Shuyi Wang, Wenhan Zhang, Yuqi Ouyang
arXiv:2609.14316v1 Announce Type: new
Abstract: Advances in image generation have made synthetic images increasingly difficult to distinguish from real photographs, raising concerns about the trustwo...
By Manni Cui, Ruiqi Liu, Zijian Yu, Hao Tan, Zibo Wei, Zian Wang, Ziheng Qin, Huijia Zhu, Weiqiang Wang, Jun Lan, Shu Wu
MLLMCLIP introduces a heterogeneous distillation framework that transfers multimodal knowledge from a generative Multimodal Large Language Model (MLLM) teacher directly into a discriminative CLIP student, eliminating the need for synthetic hard negatives. The method uses an attention-based per-layer token selection and a CKA-based distillation loss to bridge architectural differences between the two models. As a result, MLLMCLIP achieves state‑of‑the‑art compositional accuracy and improves zero‑shot classification and image‑text retrieval performance.
By Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki, Junmo Kim, Yuki Mitsufuji
arXiv:2603.24528v2 Announce Type: replace
Abstract: Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that...
By Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bart{\l}omiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer
arXiv:2512.20257v2 Announce Type: replace
Abstract: With the rise of easily accessible generative tools for creating and manipulating multimedia content, the threat of realistic synthetic alterations...
By Daniele Cardullo, Simone Teglia, Irene Amerini
Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent studies mitigate this...
arXiv:2411. 19537v2 Announce Type: replace-cross Abstract: We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content.
By Florinel-Alin Croitoru, Andrei-Iulian Hiji, Vlad Hondru, Nicolae Catalin Ristea, Paul Irofti, Marius Popescu, Cristian Rusu, Radu Tudor Ionescu, Fahad Shahbaz Khan, Mubarak Shah
arXiv:2502.14994v2 Announce Type: replace
Abstract: The rapid advancement of AI-generated video poses challenges to digital authenticity and security. Current detection methods, often trained on spec...
By Yun-Yun Tsai, Qingyuan Liu, Ruijian Zha, Victoria Li, Pengyuan Shi, Chengzhi Mao, Junfeng Yang
arXiv:2406. 09250v5 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly susceptible to sophisticated adversarial attacks, including adaptive strategies specifically designed to bypass existing defenses.
By Samar Fares, Klea Ziu, Toluwani Aremu, Nikita Durasov, Martin Tak\'a\v{c}, Pascal Fua, Ivan Laptev, Karthik Nandakumar
arXiv:2606. 19882v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) enhance the interpretability of deep learning networks by aligning the features extracted from images with natural concepts.
By Tongqing Shi, Ge Yan, Tuomas Oikarinen, Tsui-Wei Weng
We’re introducing a neural network called CLIP which efficiently learns visual concepts from natural language supervision. CLIP can be applied to any visual classification benchmark by simply providing the names of the visual categories to be recognized, similar to the “zero-shot” capabilities of GPT-2 and GPT-3.