OpenAI Blog

Disrupting a coordinated model-distillation campaign

OpenAI exposed and disrupted a coordinated campaign aimed at extracting protected model reasoning through model distillation. The incident highlighted vulnerabilities in how models can be reverse‑engineered by adversaries. In response, OpenAI is enhancing its defenses to guard against future adversarial distillation attempts.

arXiv AI
Sep 10

How to Backdoor Image Knowledge Distillation

The paper demonstrates that image knowledge distillation can be backdoored even when the teacher model is clean, by poisoning the distillation dataset with triggered and manipulated images that the teacher already classifies as a target label. The attack, effective at poisoning rates as low as 10%, uses targeted adversarial perturbations and GAN-based class transitions to embed a backdoor into the student model while preserving its performance on clean data. The study highlights that the security of knowledge distillation depends not only on the teacher but also on the integrity of the distillation data.

By Qian Ma, Chen Wu, Prasenjit Mitra, Sencun Zhu
arXiv Machine Learning
Jul 14

Reference-Based Distillation Detection in LLMs

arXiv:2607. 09692v1 Announce Type: new Abstract: Model distillation -- training on outputs from stronger third-party models -- is widely used to boost performance, but raises concerns about unfair advantages and policy violations.

By Rajat Rawat, Sizhe Chen, Akshay Anand, Michael Duan, Bob Rotsted, Sewon Min
arXiv Machine Learning
4d ago

Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

The paper introduces On-Policy Attention Self-Distillation (OPASD), a method that augments token-level supervision with solution-conditioned attention distillation for reasoning models. OPASD projects a privileged teacher’s attention onto student-visible positions, renormalizes the distribution, and aligns it with the student. Experiments on three model sizes and four math benchmarks show that OPASD improves accuracy by 4.98–8.40 percentage points, reduces generated tokens by 73.9%, cuts compute by 72.6%, and trains 1.53× faster compared to token-only distillation.

By Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin, Md Mofijul Islam
arXiv AI
Sep 2

Does Reasoning Mitigate Backdoor Attacks? A Neuro-Symbolic Perspective

The paper investigates whether neuro‑symbolic (NeSy) AI models, which combine neural perception with symbolic reasoning, can mitigate backdoor attacks. It presents a systematic evaluation comparing the NeSy framework DeepProbLog to baseline neural networks across eight backdoor settings and four reasoning tasks. Results indicate that NeSy models are generally more robust than pure neural models, but their resilience depends heavily on how strictly the reasoning process is enforced and its alignment with the attack target.

By Marco Antonio Corallo, Andrea Agiollo, Mauro Conti, Alberto Giaretta
arXiv AI
Jun 11

Diffusion-based Cumulative Adversarial Purification for Vision Language Models

arXiv:2506. 03933v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to adversarial perturbations poses a significant threat to their reliability in real-world applications.

By Jia Fu, Yongtao Wu, Yihang Chen, Kunyu Peng, Xiao Zhang, Volkan Cevher, Sepideh Pashami, Anders Holst
arXiv AI
Jun 2

Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning

arXiv:2606. 00105v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on vision-language tasks, but they may also memorize and expose sensitive or restricted knowledge, raising concerns about privacy and broader safety risks.

By Junkai Chen, Yuhao He, Junxiang You, Ruiqi Liu, Chenyu Wang, Shu Wu