arXiv Machine Learning

Position: Adversarial ML for LLMs Is Not Making Any Progress

arXiv:2502. 02260v2 Announce Type: replace Abstract: In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings.

OpenAI Blog
Feb 24, 2017

Attacking machine learning with adversarial examples

Adversarial examples are inputs to machine learning models that an attacker has intentionally designed to cause the model to make a mistake; they’re like optical illusions for machines. In this post we’ll show how adversarial examples work across different mediums, and will discuss why securing systems against them can be difficult.

arXiv AI
Sep 1

Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning

The paper introduces FAB, an attack that uses meta‑learning to embed dormant adversarial behaviors into large language models (LLMs). These behaviors remain inactive until the model is finetuned by downstream users, at which point the model can exhibit unwanted actions such as unsolicited advertising, jailbreakability, or over‑refusal. FAB is shown to be effective across multiple LLMs and resilient to various finetuning settings.

By Thibaud Gloaguen, Mark Vero, Robin Staab, Martin Vechev
arXiv Machine Learning
Sep 11

Perturbation: A simple and efficient adversarial tracer for representation learning in language models

The paper introduces Perturbation, a method that treats representations in language models as learning conduits rather than activation patterns. By fine‑tuning a model on a single adversarial example and observing how this perturbation spreads to other inputs, the approach avoids geometric assumptions and does not identify representations in untrained models. In trained models, Perturbation uncovers structured transfer across multiple linguistic scales, indicating that language models generalize along representational lines and acquire linguistic abstractions through experience.

By Joshua Rozner, Cory Shain
arXiv AI
Aug 25

A New Type of Adversarial Examples

The paper introduces a new class of adversarial examples that are markedly different from original inputs yet produce the same model output. It presents algorithms such as NI-FGSM, NI-FGM, and their momentum variants (NMI-FGSM, NMI-FGM) to generate these examples. The authors demonstrate that these adversarial examples are not confined to the vicinity of training data but are spread throughout the sample space.

By Xingyang Nie, Caoliang Zhang, Su Pan, Biao Wang, Huilin Ge, Tao Fang
arXiv Machine Learning
Jun 26

Over-parameterization and Adversarial Robustness in Neural Networks: An Overview and Empirical Analysis

arXiv:2406. 10090v3 Announce Type: replace Abstract: Thanks to their extensive capacity, over-parameterized neural networks exhibit superior predictive capabilities and generalization.

By Srishti Gupta, Zhang Chen, Luca Demetrio, Fabio Brau, Xiaoyi Feng, Zhaoqiang Xia, Antonio Emanuele Cin\`a, Maura Pintor, Luca Oneto, Ambra Demontis, Battista Biggio, Fabio Roli
arXiv AI
2d ago

UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

UniGuardian is a training‑free detector for large language models that jointly identifies prompt injection, backdoor, and adversarial attacks—collectively called Prompt Trigger Attacks (PTA). It measures how structured prompt perturbations shift the model’s output distribution and uses a single‑forward strategy to detect attacks while generating text in a shared batched forward pass. Experiments show that UniGuardian accurately and efficiently identifies trigger‑activated prompts in LLMs.

By Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao