Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents
Read the original on arXiv AI →The paper titled "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents" demonstrates how an attacker can bypass current skill‑scanning defenses by crafting malicious skills that evade both static analysis and LLM‑based semantic checks. By moving malicious payloads into natural language and distributing instructions across files, the white‑box attacker named Pretext achieves high evasion rates—up to 97% against a frozen detector and 77% against a co‑adaptive one—across three open‑source models.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.