arXiv AI

Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?

arXiv AI
Sep 7

Rethinking Indirect Prompt Injection as a Test-Time Search Problem

The paper redefines indirect prompt injection as a test-time search problem, framing it as exploration over a task-dependent attack surface shaped by the environment, user task, and injection task. It introduces an agentic attacker equipped with a search harness that conducts environment reconnaissance, structured reasoning on attack strategies, and adaptive evaluation using victim-agent feedback. Experiments across diverse tasks show that higher attacker compute leads to better vulnerability discovery, and that managing search strategies is crucial to avoid redundancy and maintain gains at larger budgets.

By Duong M. Nguyen, Joon Sik Kim, Blazej Manczak, Vaikkunth Mugunthan
arXiv AI
Sep 7

Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection

Repeat-After-Me is a black-box adaptive visual prompt injection technique that can reveal personally identifiable information or trigger malicious tool calls in both open-weight and commercial vision‑language models, achieving attack success rates above 80% on Qwen3.6‑27B and 47% on GPT‑5.5. The method works even when the benign user prompt is unrelated to the injected task and does not explicitly authorize it, and it retains significant effectiveness when transferred across models or optimized on surrogate systems. In a real‑world OpenClaw Discord deployment, a minimally injected image can overwrite TOOLS.md, enabling remote code execution and secret exfiltration.

By Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner, Arman Zharmagambetov