The paper investigates whether large, instruction‑following Vision‑Language Models (VLMs) can reliably perform zero‑shot image privacy classification. It compares three open‑source VLMs to specialized privacy models on two public benchmarks, evaluating accuracy, robustness to image degradations (compression, lighting changes, noise), inference speed, and parameter count. The findings show that while VLMs remain robust to perturbations, they are less accurate and significantly slower than smaller, purpose‑built privacy models, indicating that scaling alone does not guarantee effective privacy classification.
By Alina Elena Baia, Alessio Xompero, Andrea Cavallaro
arXiv:2608.28691v1 Announce Type: cross
Abstract: Wearable VLM pipelines promise continuous multimodal assistance from egocentric visual capture: a user asks a task-driven question about the surround...
By Zhimin Li, Pan Wang, Jingxian Chen, Yuantao Tang, Anthony Chen, Qian Lou, Jingtong Hu
Privacy risks in text-only Large Language Models (LLMs) are well studied, particularly their tendency to memorize and leak sensitive information. However, Multi-modal Large Language Models (MLLMs), which process both text and images, introduce unique privacy challenges that remain underexplored.
arXiv:2606. 09125v1 Announce Type: cross Abstract: Privacy risks in text-only Large Language Models (LLMs) are well studied, particularly their tendency to memorize and leak sensitive information.
By Tiejin Chen, Pingzhi Li, Kaixiong Zhou, Tianlong Chen, Hua Wei
arXiv:2608.21133v1 Announce Type: new
Abstract: Medical image-text data can expose protected health information (PHI) through both visible image content as well as accompanying text, creating a barri...
By Shiva Shrestha, Zongxing Xie, Chen Zhao, Liran Ma, Zhipeng Cai, Honghui Xu
The paper argues that evaluating privacy‑enhancing technologies (PETs) solely through image classification is insufficient because classification remains robust to many geometric and local perturbations. It proposes a compute‑aware multi‑task protocol that uses lightweight proxy tasks to assess PETs across various transformations, revealing that PETs with similar classification accuracy can perform very differently on other vision tasks. The study demonstrates the necessity of broader evaluation metrics beyond classification to truly gauge PET effectiveness.
By Leon Ranke, Wolfgang H\"ubner, Ronny Hug, Michael Arens, J\"urgen Beyerer
The paper introduces a new object detection approach that protects sensitive visual data in test images. It is the first to apply perceptual encryption to object detection, leveraging the embedding structure of Vision Transformers and a key-based domain adaptation to maintain accuracy comparable to unprotected models. Experiments on ViTdet confirm the method’s effectiveness in both accuracy and visual protection.
By Homare Sueyoshi, Kiyoshi Nishikawa, Hitoshi Kiya
arXiv:2610.01944v1 Announce Type: new
Abstract: Few-shot personalization enables large vision--language models (LVLMs) to learn user-specific visual concepts for applications such as personalized ret...
By Abhishek Basu, Fahad Shamshad, Karthik Nandakumar
Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are particularly acute in healthcare, where patient medical images paired with clinical reports demand rigorous privacy safeguards.
arXiv:2609.15671v1 Announce Type: cross
Abstract: Visual Question Answering (VQA) with Vision-Language Models (VLMs) is increasingly used in privacy-sensitive and bandwidth-constrained settings. Fede...
By Md Khalid Syfullah, Alvi Ataur Khalil
arXiv:2607. 08867v1 Announce Type: cross Abstract: Cloud-based deep learning enables large-scale medical image analysis but raises significant privacy concerns when sensitive patient images are outsourced for model development.
By Jason Rojas, Jiajie He, Yash Patel, Yuechun Gu, Zeyun Yu, Keke Chen
arXiv:2607. 22745v1 Announce Type: cross Abstract: Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation.
By Yi-Zhi Wang, Yichen Xiao, Linan Yue, Weibo Gao, Yichao Du, Pengfei Fang, Shimin Di, Min-Ling Zhang