arXiv AI By Yunseok Lee, Yunji Kim, Woojin Lee

Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents

Read the original on arXiv AI →

The paper introduces covert indirect prompt injection (IPI) attacks on tool‑using large language model agents, distinguishing between covert and overt successes. It defines new metrics—Covert Success Rate (CSR) and Overt Success Rate (OSR)—to capture whether users notice the injection. The authors propose ICoA, an attack that steers agents back to the user’s task after executing the injection, achieving higher CSR than existing methods on four target models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.