OpenAI Blog

Disrupting a coordinated model-distillation campaign

Read the original on OpenAI Blog →

OpenAI exposed and disrupted a coordinated campaign aimed at extracting protected model reasoning through model distillation. The incident highlighted vulnerabilities in how models can be reverse‑engineered by adversaries. In response, OpenAI is enhancing its defenses to guard against future adversarial distillation attempts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

arXiv AI
Sep 10

How to Backdoor Image Knowledge Distillation

The paper demonstrates that image knowledge distillation can be backdoored even when the teacher model is clean, by poisoning the distillation dataset with triggered and manipulated images that the teacher already classifies as a target label. The attack, effective at poisoning rates as low as 10%, uses targeted adversarial perturbations and GAN-based class transitions to embed a backdoor into the student model while preserving its performance on clean data. The study highlights that the security of knowledge distillation depends not only on the teacher but also on the integrity of the distillation data.

By Qian Ma, Chen Wu, Prasenjit Mitra, Sencun Zhu