arXiv AI By Thomas Thebaud, Sonal Joshi, Henry Li, Martin Sustek, Jesus Villalba, Sanjeev Khudanpur, Najim Dehak

Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System

Read the original on arXiv AI →

arXiv:2606. 28953v1 Announce Type: cross Abstract: Poisoning attacks entail attackers intentionally tampering with training data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples

SAGE is a defense against clean‑label data poisoning that relies on a very small set of verified examples—both clean and poisoned—rather than a large clean set. It trains a generic feature extractor on a separate dataset and then uses a non‑parametric, similarity‑weighted prediction to flag poisoned training examples. Experiments on standard benchmarks show that even a handful of verified poisoned examples give a substantial advantage, and that the distribution of verified clean examples across classes is more important than their sheer number.

By Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik, Aron Laszka
arXiv AI
2d ago

UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

UniGuardian is a training‑free detector for large language models that jointly identifies prompt injection, backdoor, and adversarial attacks—collectively called Prompt Trigger Attacks (PTA). It measures how structured prompt perturbations shift the model’s output distribution and uses a single‑forward strategy to detect attacks while generating text in a shared batched forward pass. Experiments show that UniGuardian accurately and efficiently identifies trigger‑activated prompts in LLMs.

By Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao