arXiv AI

ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models

The paper introduces the ASCII Attack, a single‑turn, black‑box method that embeds a harmful request within ASCII art and presents it as artwork to a large language model. By framing the request as artistic critique, the model can provide operational details that a plain request would normally be refused. Experiments across eleven models and eight harm topics show that the attack succeeds in 62% of cases versus 42% for direct controls, with the most vulnerable model achieving a 93% success rate.

arXiv Computation and Language
Aug 27

When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models

The paper introduces Paraesthesia, a dynamic backdoor attack that uses emotionally styled inputs as triggers for large language models. By mapping target emotions into a valence–arousal space and rewriting a small subset of clean samples, the attack achieves over 98% success while minimally affecting clean performance. Experiments on four major LLMs show that the trigger cannot be fully explained by token-level cues and remains robust against several filtering and mitigation techniques.

By Ziyu Liu, Tao Li, Tao Yang, Tianjie Ni, Xiaolong Lan, Wengang Ma, Junjiang He
arXiv AI
Sep 7

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

The paper investigates how large language models balance helpfulness and safety by refusing harmful queries while responding to benign ones. It decomposes safety-tuning responses into a boilerplate refusal statement and a rationale, finding that the statement causes false refusals by relying on superficial cues. Training on rationales alone reduces false refusals without compromising safety performance, suggesting that fine‑grained safety supervision is essential for better alignment.

By Minji Kim, Hyounghun Kim
arXiv AI
Aug 11

Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks

arXiv:2607. 26574v2 Announce Type: replace-cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap.

By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Yi Feng, Xiao Luo, Zijian Xiao, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv AI
Aug 24

Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification

The paper identifies a vulnerability in large language models where harmful intent can be hidden within benign narratives, a phenomenon termed Semantic Camouflage. By examining latent activation patterns across several small language model families, the authors discover an "Intent Horizon"—a layer depth where harmful intent representations collapse. They propose Latent Intent Verification (LIV), a lightweight probing defense that detects harmful intent in early layers and outperforms existing guardrails on the PKU-SafeRLHF dataset.

By Md. Hasib Ur Rahman
arXiv AI
Sep 4

CASCADE: A Component Ablation and Corpus Audit of a Layered Local Defense for MCP-Based Systems

The paper evaluates CASCADE, a fully local layered defense for Model Context Protocol (MCP)-based systems, by conducting a component ablation and corpus audit on a fixed 5,000-sample dataset. It demonstrates that the choice of aggregation convention heavily influences reported metrics, that detection performance varies with provenance, and that the released configuration does not fully disclose the operating point. The study also shows that a local review model invoked for a third of requests does not alter classification outcomes, highlighting the importance of reproducibility and transparency in defense evaluations.

By \.Ipek Abas{\i}kele\c{s} Turgut, Edip G\"um\"u\c{s}