arXiv AI By Neha Nagaraja, Amisha Bagari, Hayretdin Bahsi

When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems

Read the original on arXiv AI →

arXiv:2608. 00747v2 Announce Type: replace-cross Abstract: Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration exposes them to prompt injection attacks that can lead to unsafe decisions and physical harm.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 7

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots

arXiv:2608. 05715v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding.

By S. M . Bhagya P. Samarakoon, M. A. Viraj J. Muthugala, W. K. R. Sachinthana, Mohan Rajesh Elara
arXiv AI
Sep 16

Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks

The paper proposes universal, tool‑based defenses for large language model agents that use external tools, addressing four types of adversarial attacks: direct and indirect prompt injection, memory poisoning, and backdoor attacks. Two main defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to remove suspicious tools, and Normal Tool Recalling, which restores the agent’s original toolset before planning. The authors also add prompt‑based defenses such as Chain‑of‑Thought prompting and self‑reflection, and demonstrate that these methods dramatically lower attack success rates—often to 0%—across multiple open‑source and proprietary LLMs while maintaining or improving task performance.

By Xiaoyan Li, Yunli Wang
arXiv AI
Sep 2

Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents

The paper introduces covert indirect prompt injection (IPI) attacks on tool‑using large language model agents, distinguishing between covert and overt successes. It defines new metrics—Covert Success Rate (CSR) and Overt Success Rate (OSR)—to capture whether users notice the injection. The authors propose ICoA, an attack that steers agents back to the user’s task after executing the injection, achieving higher CSR than existing methods on four target models.

By Yunseok Lee, Yunji Kim, Woojin Lee