arXiv Machine Learning By Kristina Nikoli\'c, Egor Zverev, Javier Rando, Matthew Jagielski, Edoardo Debenedetti, Florian Tram\`er

Untrusted Content Masking for Web Agents with Security Guarantees

Read the original on arXiv Machine Learning →

arXiv:2607. 05277v1 Announce Type: cross Abstract: Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 16

MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks

arXiv:2602. 09222v2 Announce Type: replace-cross Abstract: Large language model (LLM) based web agents are increasingly deployed to automate complex online tasks by directly interacting with web sites and performing actions on users' behalf.

By Georgios Syros, Evan Rose, Brian Grinstead, Christoph Kerschbaumer, William Robertson, Cristina Nita-Rotaru, Alina Oprea