UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
Read the original on arXiv AI →UniGuardian is a training‑free detector for large language models that jointly identifies prompt injection, backdoor, and adversarial attacks—collectively called Prompt Trigger Attacks (PTA). It measures how structured prompt perturbations shift the model’s output distribution and uses a single‑forward strategy to detect attacks while generating text in a shared batched forward pass. Experiments show that UniGuardian accurately and efficiently identifies trigger‑activated prompts in LLMs.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.