arXiv:2607. 01702v1 Announce Type: cross Abstract: Recently, speech classification methods have gained widespread adoption in intelligent gadgets.
By Yueming Huang, Wenhan Yao, Fen Xiao, Xiarun Chen, Weiping Wen
Recently, speech classification methods have gained widespread adoption in intelligent gadgets. Current study indicates that backdoor attacks provide a substantial security concern to these models, underscoring the pressing necessity to investigate additional potential attack techniques to expose and prevent such risks.
SAGE is a defense against clean‑label data poisoning that relies on a very small set of verified examples—both clean and poisoned—rather than a large clean set. It trains a generic feature extractor on a separate dataset and then uses a non‑parametric, similarity‑weighted prediction to flag poisoned training examples. Experiments on standard benchmarks show that even a handful of verified poisoned examples give a substantial advantage, and that the distribution of verified clean examples across classes is more important than their sheer number.
By Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik, Aron Laszka
arXiv:2607. 23394v1 Announce Type: new Abstract: Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly.
By Adhyyan Narang, Artin Tajdini, Claire Zhang, Jamie Morgenstern
UniGuardian is a training‑free detector for large language models that jointly identifies prompt injection, backdoor, and adversarial attacks—collectively called Prompt Trigger Attacks (PTA). It measures how structured prompt perturbations shift the model’s output distribution and uses a single‑forward strategy to detect attacks while generating text in a shared batched forward pass. Experiments show that UniGuardian accurately and efficiently identifies trigger‑activated prompts in LLMs.
By Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao
arXiv:2607. 15697v1 Announce Type: cross Abstract: Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data.
By Jinwen Xin, Xixiang Lv
Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or training-time mitigation, but face two key limitations.
arXiv:2405.18540v3 Announce Type: replace-cross
Abstract: Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of larg...
By Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre, Juho Lee, Sung Ju Hwang, Kenji Kawaguchi, Gauthier Gidel, Yoshua Bengio, Esmeralda S. Whitammer, Moksh Jain
The paper investigates poisoning-based backdoor attacks on Speech Emotion Recognition (SER) systems that use self‑supervised acoustic representations. It introduces a stealthy, low‑energy acoustic trigger that can be embedded imperceptibly into both natural and synthetic speech, enabling scalable poisoning. Experiments show high attack success rates with low poisoning ratios, cross‑model transferability, and a particular vulnerability of self‑supervised representations, highlighting the lowered barrier to effective backdoor attacks via TTS technology.
By Yongbin Huang, Xihao Xie, Jia Zhang
Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses target backdoors one at a time and typically require knowledge of the trigger, leaving the defender at a structural disadvantage when unknown backdoors may exist in a model.
arXiv:2605. 26595v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison.
By Zedian Shao, Charles Fleming, Teodora Baluta
SynGhost is a novel task‑agnostic backdoor attack that injects invisible syntactic backdoors into pre‑training corpora of language models. It uses an entropy‑based poisoning filter, contrastive learning to select optimal targets, and an awareness module to reduce interference between backdoors, thereby preserving the model’s pre‑training performance. Experiments demonstrate that SynGhost can transfer to multiple downstream tasks and withstand several defense mechanisms such as perplexity checks, fine‑pruning, and the maxEntropy filter.
By Pengzhou Cheng, Wei Du, Zongru Wu, Fengwei Zhang, Libo Chen, Zhuosheng Zhang, Gongshen Liu