arXiv AI

"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild

arXiv:2602. 06547v4 Announce Type: replace-cross Abstract: LLM-based coding agents increasingly rely on third-party extensions called skills, which bundle natural language instructions and helper scripts that execute with full user privileges.

arXiv Computation and Language
Aug 25

SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents

The paper "SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents" investigates how agent skills—task‑specific instructions, scripts, and resources—can be exploited to create a trusted instruction channel that enables token amplification attacks. It introduces a two‑phase framework, SkillBloat, which first screens a library of attack‑type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM‑guided full‑document skill rewriting. Evaluated on a real‑world skill benchmark, SkillBloat achieves an average best amplification of 5.4184×–10.1455× across multiple coding‑agent target configurations, and an ablation study shows that the second‑stage refinement consistently improves performance over the initial screening alone.

By Yuanjin Zheng, Jingbang Chen
arXiv AI
Jun 16

MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks

arXiv:2602. 09222v2 Announce Type: replace-cross Abstract: Large language model (LLM) based web agents are increasingly deployed to automate complex online tasks by directly interacting with web sites and performing actions on users' behalf.

By Georgios Syros, Evan Rose, Brian Grinstead, Christoph Kerschbaumer, William Robertson, Cristina Nita-Rotaru, Alina Oprea
arXiv AI
5d ago

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

The paper introduces skill cascading attacks, where a malicious goal is spread across multiple seemingly benign skills, causing harmful outcomes when combined. It presents SkillCascade, an automated red‑teaming framework, and releases SkillCascade‑Bench, a benchmark of 213 validated cascading test cases across various agent systems and domains. Experiments show that these cascaded interactions reliably induce harmful behaviors while evading existing per‑skill scanners and runtime monitors, revealing a gap between component‑level integrity and system‑level safety.

By Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu