arXiv AI By Yinghan Hou, Zongyou Yang

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills

Read the original on arXiv AI →

arXiv:2604. 06550v3 Announce Type: replace-cross Abstract: Agent skills combine natural-language instructions with executable code while inheriting an agent's filesystem, credential, and network access.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

TRUSS is a framework that generates and verifies automated agent skills, ensuring they are both functionally effective and safe. It evaluates candidate skills against source evidence and nine safety properties, then tests them in a controlled environment to capture execution traces and identify failures. The system iteratively refines skills based on these results, achieving high precision in vulnerability detection and significantly improving task performance and security rates.

By Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang
arXiv AI
3d ago

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

The paper titled "Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents" demonstrates how an attacker can bypass current skill‑scanning defenses by crafting malicious skills that evade both static analysis and LLM‑based semantic checks. By moving malicious payloads into natural language and distributing instructions across files, the white‑box attacker named Pretext achieves high evasion rates—up to 97% against a frozen detector and 77% against a co‑adaptive one—across three open‑source models.

By Tobias Kaisar, Aritra Dhar