Hugging Face Trending Papers

FIDA: Feature Instability-Driven Attack on Self-Supervised Facial Representation

FIDA (Feature Instability-Driven Attack) is a new backdoor attack framework targeting self-supervised facial representation models. It employs subtle semantic triggers and a novel Feature Instability Loss to make the encoder more sensitive to perturbations, thereby avoiding the rigid feature patterns seen in earlier attacks. Experiments demonstrate that FIDA achieves high attack success while largely preserving normal model performance.

arXiv Computer Vision
Aug 28

FIDA: Feature Instability-Driven Attack on Self-Supervised Facial Representation

FIDA (Feature Instability-Driven Attack) is a new backdoor attack framework targeting self‑supervised facial representation models. It employs subtle semantic triggers and a novel Feature Instability Loss that trains the encoder to heighten sensitivity of triggered features along perturbation directions, thereby avoiding the rigid feature patterns seen in prior attacks. Experiments demonstrate that FIDA achieves high attack success while largely preserving benign utility, exposing a significant threat to real‑world facial analysis applications.

By Zhiyang Chen, Changchun Yin, Huiqin Yang, Liming Fang
Hugging Face Trending Papers
Jul 20

DecoyFace: Beyond Obfuscation via Controllable and Imperceptible Identity Misdirection for Privacy-Preserving Face Recognition

Split face recognition reduces client-side computation but exposes intermediate features to feature inversion attacks and unauthorized analysis by honest-but-curious (HBC) servers. Existing privacy-preserving face recognition methods mainly aim to resist unauthorized reconstruction, typically producing features whose inversion yields visibly degraded results, which may reveal the existence of protection and motivate adaptive attacks.

arXiv Computer Vision
Aug 27

Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models

The paper introduces Learning to Detect (LoD), a framework for identifying unseen jailbreak attacks in Large Vision‑Language Models without relying on attack data or hand‑crafted heuristics. LoD extracts layer‑wise safety representations via Multi‑modal Safety Concept Activation Vectors and compresses them into a one‑dimensional anomaly score using a Safety Pattern Auto‑Encoder. Experiments show that LoD achieves state‑of‑the‑art AUROC across diverse unseen attacks on multiple LVLMs while improving efficiency.

By Shuang Liang, Zhihao Xu, Jiaqi Weng, Jialing Tao, Hui Xue, Xiting Wang