arXiv Machine Learning By Laiqiao Qin, Tianqing Zhu, Longxiang Gao, Wanlei Zhou

BASIS: Breach-Aware Selective Prompt Injection Shielding with Prefill Attention Probes

Read the original on arXiv Machine Learning →

arXiv:2608. 08027v1 Announce Type: cross Abstract: Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
6d ago

From ASR to ASP: Evaluating Prompt Attack Vulnerabilities Against Open-Source LLMs

The paper investigates prompt injection attacks on 14 open‑source and 3 closed‑source large language models (LLMs), introducing a new metric called Attack Success Probability (ASP) that accounts for uncertainty in model responses. It demonstrates that a simple hypnotism attack can trigger objectionable behavior in models such as StableLM2, Mistral, Openchat, and Vicuna, achieving roughly 90% ASP. The study highlights that moderately well‑known LLMs are particularly vulnerable, underscoring the importance of public awareness and effective mitigation strategies.

By Jiawen Wang, Pritha Gupta, Eyke H\"ullermeier, Xiaoxue Gao, Nancy F. Chen
arXiv Computation and Language
6d ago

Prompt Injection Detection for Email Agents Through Attack Chain Modeling

The paper introduces a prompt‑injection detection framework for email assistants that models attacks as a chain of stages. It combines a text detector, stage‑specific verifiers, rule‑based risk signals, user intent consistency checks, and a logistic decision policy. Experiments on five benchmarks show the framework outperforms pretrained detectors, achieving a mean F1 of 0.406 versus 0.216, and demonstrate that training on benign emails resembling attacks reduces false alarms.

By Ahmad Hashmi, Dhyey Patel, Yunting Yin
arXiv Machine Learning
1d ago

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

The paper introduces NEEDLE, a training‑free technique for removing backdoors from large language models. After a trigger is identified, NEEDLE estimates a backdoor direction and a refusal subspace using activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preserving refusal‑related representations. The method requires no clean reference model or original poisoned data and achieves the lowest attack success rate and minimal impact on model performance across multiple model families and attack types.

By Minoo Kim, Vasileios Lampos, George Drayson