SODA: Semi On-Policy Black-Box Distillation for Large Language Models
arXiv:2604. 03873v4 Announce Type: replace Abstract: Black-box knowledge distillation for large language models presents a strict trade-off.
arXiv:2607. 27737v1 Announce Type: new Abstract: Deep neural networks (DNNs) have achieved remarkable success in classical machine learning problems.
arXiv:2604. 03873v4 Announce Type: replace Abstract: Black-box knowledge distillation for large language models presents a strict trade-off.
The paper demonstrates that image knowledge distillation can be backdoored even when the teacher model is clean, by poisoning the distillation dataset with triggered and manipulated images that the teacher already classifies as a target label. The attack, effective at poisoning rates as low as 10%, uses targeted adversarial perturbations and GAN-based class transitions to embed a backdoor into the student model while preserving its performance on clean data. The study highlights that the security of knowledge distillation depends not only on the teacher but also on the integrity of the distillation data.
arXiv:2412. 08394v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are vulnerable to adversarial samples crafted by adding imperceptible perturbations to clean data, potentially leading to incorrect and dangerous predictions.
arXiv:2606. 01437v1 Announce Type: cross Abstract: Deep Neural Networks (DNNs) are highly susceptible to adversarial perturbations, leading to extensive research on robustness for safety-critical applications.
arXiv:2606. 31653v1 Announce Type: cross Abstract: Certified training aims to produce models whose predictions can be formally verified against adversarial perturbations, typically by optimising upper bounds on the worst-case loss over an allowed perturbation set.
arXiv:2606. 01746v1 Announce Type: cross Abstract: Modern neural networks are highly susceptible to adversarial perturbations.
arXiv:2511. 13749v2 Announce Type: replace Abstract: Deep neural networks are known to be vulnerable to adversarial perturbations, which are small, carefully crafted inputs that lead to incorrect predictions.
arXiv:2606. 00105v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on vision-language tasks, but they may also memorize and expose sensitive or restricted knowledge, raising concerns about privacy and broader safety risks.
arXiv:2607. 04145v1 Announce Type: new Abstract: Adversarial attacks guide and provide additional training and test data for both adversarial training and adversarial robustness validation, and expose the 'piecewise linearity' of deep learning based models.
arXiv:2606. 02267v1 Announce Type: new Abstract: The vulnerability of deep neural networks to adversarial examples poses a significant challenge for real-world deployment.
arXiv:2606. 27784v1 Announce Type: cross Abstract: The existence of adversarial attacks is often attributed to the presence of non-robust features in neural networks.
arXiv:2607. 03075v1 Announce Type: new Abstract: Safety-critical applications require classifiers that are both robust and reliable.