Vision Language Model Helps Private Information De-Identification in Vision Data
arXiv:2606. 09132v1 Announce Type: new Abstract: Visual Language Models (VLMs) have gained significant popularity due to their remarkable ability.
The paper introduces a new object detection approach that protects sensitive visual data in test images. It is the first to apply perceptual encryption to object detection, leveraging the embedding structure of Vision Transformers and a key-based domain adaptation to maintain accuracy comparable to unprotected models. Experiments on ViTdet confirm the method’s effectiveness in both accuracy and visual protection.
arXiv:2606. 09132v1 Announce Type: new Abstract: Visual Language Models (VLMs) have gained significant popularity due to their remarkable ability.
The paper argues that evaluating privacy‑enhancing technologies (PETs) solely through image classification is insufficient because classification remains robust to many geometric and local perturbations. It proposes a compute‑aware multi‑task protocol that uses lightweight proxy tasks to assess PETs across various transformations, revealing that PETs with similar classification accuracy can perform very differently on other vision tasks. The study demonstrates the necessity of broader evaluation metrics beyond classification to truly gauge PET effectiveness.
The paper investigates whether large, instruction‑following Vision‑Language Models (VLMs) can reliably perform zero‑shot image privacy classification. It compares three open‑source VLMs to specialized privacy models on two public benchmarks, evaluating accuracy, robustness to image degradations (compression, lighting changes, noise), inference speed, and parameter count. The findings show that while VLMs remain robust to perturbations, they are less accurate and significantly slower than smaller, purpose‑built privacy models, indicating that scaling alone does not guarantee effective privacy classification.
arXiv:2607. 08867v1 Announce Type: cross Abstract: Cloud-based deep learning enables large-scale medical image analysis but raises significant privacy concerns when sensitive patient images are outsourced for model development.
The unprecedented growth of computer vision applications, such as surveillance systems and social media, raises security and visual privacy concerns, especially when data is stored on cloud servers. Image obfuscation offers a way to preserve visual privacy while maintaining an adequate level of usability; thus, it has been a topic of great interest in recent years.
The paper investigates federated adversarial training (AT) for vision transformers, a topic not previously explored in federated learning (FL). It evaluates various transformer architectures and aggregation strategies, and introduces FedWAvg, an extension of FedAvg that weights client updates based on similarity of their last-layer representations. Experiments demonstrate that FedWAvg yields higher robust accuracy than existing aggregation methods in non‑IID settings.
arXiv:2608.28691v1 Announce Type: cross Abstract: Wearable VLM pipelines promise continuous multimodal assistance from egocentric visual capture: a user asks a task-driven question about the surround...
arXiv:2607. 08402v1 Announce Type: cross Abstract: Large-scale and diverse datasets are needed to train AI models to take real-time decisions for autonomous vehicles (AVs), an intelligent transportation system (ITS) application.
arXiv:2607. 14932v1 Announce Type: cross Abstract: Synthetic face datasets have become effective enough to train face recognition models with accuracy rivaling that of models trained on real photographs.
arXiv:2608.23012v1 Announce Type: new Abstract: Image matching is a core component of applications such as Simultaneous Localization and Mapping (SLAM), Visual Localization, and Structure from Motion...
Synthetic face datasets have become effective enough to train face recognition models with accuracy rivaling that of models trained on real photographs. This progress sidesteps the ethical and legal burdens of collecting real biometric data, yet evaluation has not kept pace.
Image matching is a core component of applications such as Simultaneous Localization and Mapping (SLAM), Visual Localization, and Structure from Motion (SfM). However, the local image features central...