arXiv AI

D-ADD: An Effective Plug-In for Defending Against Model Stealing

The paper introduces D-ADD, a plug‑in defense for image classification models that protects against model‑stealing attacks. It uses a non‑parametric detector called Account‑aware Distribution Discrepancy (ADD) to identify malicious queries by modeling each class as a multivariate normal distribution and computing weighted distribution discrepancies. With an enhanced version ADD$^+$ that handles domain shifts and combined with random‑based prediction poisoning, D-ADD offers strong protection while minimally affecting benign users in both soft‑ and hard‑label settings.

arXiv Machine Learning
Jun 8

ADAGE: Active Defenses Against GNN Extraction

arXiv:2503. 00065v4 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) achieve high performance in various real-world applications, such as drug discovery, traffic states prediction, and recommendation systems.

By Jing Xu, Franziska Boenisch, Adam Dziedzic
arXiv AI
Sep 1

Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning

The paper introduces FAB, an attack that uses meta‑learning to embed dormant adversarial behaviors into large language models (LLMs). These behaviors remain inactive until the model is finetuned by downstream users, at which point the model can exhibit unwanted actions such as unsolicited advertising, jailbreakability, or over‑refusal. FAB is shown to be effective across multiple LLMs and resilient to various finetuning settings.

By Thibaud Gloaguen, Mark Vero, Robin Staab, Martin Vechev