Model Distillation in the API
Fine-tune a cost-efficient model with the outputs of a large frontier model–all on the OpenAI platform
OpenAI exposed and disrupted a coordinated campaign aimed at extracting protected model reasoning through model distillation. The incident highlighted vulnerabilities in how models can be reverse‑engineered by adversaries. In response, OpenAI is enhancing its defenses to guard against future adversarial distillation attempts.
Fine-tune a cost-efficient model with the outputs of a large frontier model–all on the OpenAI platform
arXiv:2604. 03873v4 Announce Type: replace Abstract: Black-box knowledge distillation for large language models presents a strict trade-off.
The paper demonstrates that image knowledge distillation can be backdoored even when the teacher model is clean, by poisoning the distillation dataset with triggered and manipulated images that the teacher already classifies as a target label. The attack, effective at poisoning rates as low as 10%, uses targeted adversarial perturbations and GAN-based class transitions to embed a backdoor into the student model while preserving its performance on clean data. The study highlights that the security of knowledge distillation depends not only on the teacher but also on the integrity of the distillation data.
arXiv:2606. 09091v1 Announce Type: new Abstract: On-policy distillation (OPD) has recently emerged as an important post-training paradigm.
arXiv:2604. 05634v2 Announce Type: replace Abstract: Machine unlearning (MU) has become a critical technique for GenAI models' safe and compliant operation.
arXiv:2607. 04751v1 Announce Type: cross Abstract: Big goals are hard to achieve all at once; breaking them into small steps is wiser.
arXiv:2607. 09692v1 Announce Type: new Abstract: Model distillation -- training on outputs from stronger third-party models -- is widely used to boost performance, but raises concerns about unfair advantages and policy violations.
The paper introduces On-Policy Attention Self-Distillation (OPASD), a method that augments token-level supervision with solution-conditioned attention distillation for reasoning models. OPASD projects a privileged teacher’s attention onto student-visible positions, renormalizes the distribution, and aligns it with the student. Experiments on three model sizes and four math benchmarks show that OPASD improves accuracy by 4.98–8.40 percentage points, reduces generated tokens by 73.9%, cuts compute by 72.6%, and trains 1.53× faster compared to token-only distillation.
The paper investigates whether neuro‑symbolic (NeSy) AI models, which combine neural perception with symbolic reasoning, can mitigate backdoor attacks. It presents a systematic evaluation comparing the NeSy framework DeepProbLog to baseline neural networks across eight backdoor settings and four reasoning tasks. Results indicate that NeSy models are generally more robust than pure neural models, but their resilience depends heavily on how strictly the reasoning process is enforced and its alignment with the attack target.
arXiv:2506. 03933v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to adversarial perturbations poses a significant threat to their reliability in real-world applications.
arXiv:2607. 27737v1 Announce Type: new Abstract: Deep neural networks (DNNs) have achieved remarkable success in classical machine learning problems.
arXiv:2606. 00105v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on vision-language tasks, but they may also memorize and expose sensitive or restricted knowledge, raising concerns about privacy and broader safety risks.