arXiv AI By Mudit Sinha, Sanika Chavan

Hiding in Plain Floats: Steganographic Carriers for Indirect Prompt and Content Injection

Read the original on arXiv AI →

arXiv:2606. 08403v1 Announce Type: cross Abstract: Text-centered prompt-injection defenses assume that the malicious signal is visible in one of the inspected text views.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection

Repeat-After-Me is a black-box adaptive visual prompt injection technique that can reveal personally identifiable information or trigger malicious tool calls in both open-weight and commercial vision‑language models, achieving attack success rates above 80% on Qwen3.6‑27B and 47% on GPT‑5.5. The method works even when the benign user prompt is unrelated to the injected task and does not explicitly authorize it, and it retains significant effectiveness when transferred across models or optimized on surrogate systems. In a real‑world OpenClaw Discord deployment, a minimally injected image can overwrite TOOLS.md, enabling remote code execution and secret exfiltration.

By Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner, Arman Zharmagambetov
arXiv AI
Sep 25

Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation

The paper critiques current encoded‑prompt safety benchmarks that focus only on harmful requests, showing that such tests can misrepresent a model’s safety. By evaluating the benign arm under the same encoding, the authors reveal a substantial drop in the harm gap—sometimes to zero—indicating that the encoding masks true safety deficiencies. Across multiple large models and training pipelines, they document that the encoding can either hide or falsely inflate safety metrics, and they identify twelve specific instrument defects that contribute to these misleading results.

By Haoyu Zhang, Haowen Xu, Xiao Luo, Hanwen Liu, Yang Chen, Zijian Xiao, Yi Feng, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv AI
4d ago

Render Before Reading: Visual Rendering as a Prompt Injection Defense

The paper investigates how multimodal large language models are more susceptible to prompt injection when adversarial instructions are presented as text rather than as non-textual inputs like images. It proposes a training‑free defense that renders untrusted payloads into typographic images (or audio) before they reach the model, a method called Pictionary. Experiments on ten models and two benchmarks show that this approach significantly lowers attack success rates while maintaining normal functionality, and that fine‑tuning on image‑rendered instructions can further reduce the modality gap.

By Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran, Florian Tram\`{e}r