Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper titled "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents" demonstrates how an attacker can bypass current skill‑scanning defenses by crafting malicious skills that evade both static analysis and LLM‑based semantic checks. By moving malicious payloads into natural language and distributing instructions across files, the white‑box attacker named Pretext achieves high evasion rates—up to 97% against a frozen detector and 77% against a co‑adaptive one—across three open‑source models.
arXiv:2606. 00925v1 Announce Type: cross Abstract: Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime.
arXiv:2608. 08468v1 Announce Type: cross Abstract: Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security properties remain under-explored.
arXiv:2609.36879v1 Announce Type: cross Abstract: As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent...
arXiv:2604. 06550v3 Announce Type: replace-cross Abstract: Agent skills combine natural-language instructions with executable code while inheriting an agent's filesystem, credential, and network access.
arXiv:2609.39065v1 Announce Type: cross Abstract: LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabi...