arXiv Machine Learning By Qin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang, Meisam Mohammady, Doowon Kim, Yuan Hong

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks

Read the original on arXiv Machine Learning →

arXiv:2606. 09700v1 Announce Type: cross Abstract: Large language model (LLM)-powered content moderation systems have become a critical defense against harmful online content.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
4d ago

Render Before Reading: Visual Rendering as a Prompt Injection Defense

The paper investigates how multimodal large language models are more susceptible to prompt injection when adversarial instructions are presented as text rather than as non-textual inputs like images. It proposes a training‑free defense that renders untrusted payloads into typographic images (or audio) before they reach the model, a method called Pictionary. Experiments on ten models and two benchmarks show that this approach significantly lowers attack success rates while maintaining normal functionality, and that fine‑tuning on image‑rendered instructions can further reduce the modality gap.

By Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran, Florian Tram\`{e}r
arXiv Computation and Language
Aug 31

EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion

EvoHarmBench is a dynamic adversarial evaluation framework that simulates how users iteratively modify harmful content to evade moderation. It uses an optimization loop that evolves evasion strategies at the semantic-cluster level while maintaining human readability, and tests 229 semantic sub-clusters across five violation categories derived from 5,002 real-world adversarial samples. The study shows that even state‑of‑the‑art LLM‑based moderators can be bypassed with an 80.3% success rate after twelve iterations, highlighting significant vulnerabilities in current systems.

By Ruijie Jian, Benlei Cui, Ting Ma, Haidong Ding, Kangwei Liu, Ziwen Xu, Longtao Huang, Hui Xue, Ziqiang Zhu, Junjie Li, Haiwen Hong
Hugging Face Trending Papers
Jul 7

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploit them to synthesize illicit content. Existing alignment paradigms, from input sanitization to structural feature pruning, are largely organized around unsafe concepts explicitly exposed during filtering, editing, or localization.