arXiv AI

Is My Vision-Language Data in Your AI? Membership Inference Test (MINT) Demo 2

arXiv:2606. 14748v1 Announce Type: cross Abstract: We present the Membership Inference Test (MINT) Demo 2, a framework designed to improve transparency in machine learning training processes.

arXiv Machine Learning
Aug 19

MemCatalyst: Amplifying Data Auditing on Vision-Language Models via Data Poisoning

MemCatalyst is a set of data poisoning tools designed to improve data auditing for Vision‑Language Models (VLMs). It introduces two poisoning strategies—Poisoning Text and Poisoning Image—to force VLMs to over‑learn inconsistencies between image features and textual semantics, thereby increasing their vulnerability to membership inference attacks. Experiments on two prominent VLMs show that MemCatalyst significantly boosts MI AUC scores with a small number of poisoned samples while barely affecting overall model performance.

By Xukun Luan, Jinyan Liu, Yuhui Gong, Yuanguo Bi, Bing Hu, Xuesong Li, Di Wang
arXiv AI
Jul 7

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

arXiv:2607. 04163v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering.

By Kai Tang, Jinhao You, Bohua Zhang, Yichen Guo, Yiding Sun, Dongxu Zhang, Chenxi Li, Xiande Huang, Shanghang Zhang
arXiv AI
Jun 4

DetectZoo: A Unified Toolkit for AI-Generated Content Detection Across Text, Audio, and Image Modalities

arXiv:2606. 04205v1 Announce Type: cross Abstract: The growing popularity and capacity of generative models have eroded the distinction between human and machine-generated content, motivating a growing body of work on detection across text, images, and audio.

By Sajad Ebrahimi, Nima Jamali, Bardia Shirsalimian, Kelly McConvey, Wentao Zhang, Jalehsadat Mahdavimoghaddam, Maksym Taranukhin, Maura Grossman, Vered Shwartz, Yuntian Deng, Ebrahim Bagheri
arXiv AI
Sep 3

Swin Meets EfficientNet: Lightweight Architectures for GAN-Based Face Forensics

The paper presents lightweight architectures for detecting GAN-generated synthetic faces, comparing a compact Swin Transformer, pre‑trained Swin‑Tiny and Swin‑Small models, and a hybrid EfficientNet‑B0 + Swin Transformer. Using the 140K Real and Fake Faces dataset, the hybrid model achieved 99% accuracy and 99.44% recall on 5,000 test images, outperforming both pure Swin variants and a CNN‑only baseline. The study demonstrates that combining hierarchical CNN features with shifted‑window self‑attention yields an efficient, computationally lightweight detection method.

By Sejuti Basu, Ashima Sood, Vijay Kumar, Sahil Sharma