arXiv AI

What Does It Mean to Break a Distillation Defense?

arXiv:2606. 25059v2 Announce Type: replace-cross Abstract: Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a student on its outputs.

arXiv AI
Sep 10

How to Backdoor Image Knowledge Distillation

The paper demonstrates that image knowledge distillation can be backdoored even when the teacher model is clean, by poisoning the distillation dataset with triggered and manipulated images that the teacher already classifies as a target label. The attack, effective at poisoning rates as low as 10%, uses targeted adversarial perturbations and GAN-based class transitions to embed a backdoor into the student model while preserving its performance on clean data. The study highlights that the security of knowledge distillation depends not only on the teacher but also on the integrity of the distillation data.

By Qian Ma, Chen Wu, Prasenjit Mitra, Sencun Zhu
arXiv AI
Jul 7

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

arXiv:2510. 15476v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as interfaces to information, code, and real-world services, making prompt-level security failures a practical concern.

By Hanbin Hong, Shuang Wu, Shuya Feng, Nima Naderloui, Shenao Yan, Jingyu Zhang, Ali Arastehfard, Heqing Huang, Yuan Hong
arXiv Machine Learning
Sep 15

Are Targeted Data Poisoning Attacks as Effective as We Think?

arXiv:2509.06896v3 Announce Type: replace Abstract: Targeted data poisoning attacks manipulate model predictions on specific test samples by injecting malicious data into training. Yet existing evalu...

By William Xu, Chenyu Zhang, Yihan Wang, Matthew Y. R. Yang, Zuoqiu Liu, Yaoliang Yu, Gautam Kamath, Yiwei Lu