Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems
arXiv:2606. 00925v1 Announce Type: cross Abstract: Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime.
arXiv:2603. 16572v2 Announce Type: replace-cross Abstract: Agent skills extend local AI agents, such as Claude Code and OpenClaw, with additional functionality.
arXiv:2606. 00925v1 Announce Type: cross Abstract: Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime.
arXiv:2602. 06547v4 Announce Type: replace-cross Abstract: LLM-based coding agents increasingly rely on third-party extensions called skills, which bundle natural language instructions and helper scripts that execute with full user privileges.
arXiv:2602. 06547v3 Announce Type: replace-cross Abstract: LLM-based coding agents increasingly rely on third-party extensions called skills, which bundle natural language instructions and helper scripts that execute with full user privileges.
Coding agents are increasingly used for software engineering tasks, including bootstrapping projects from third-party repositories whose integrity cannot be assumed. Prior work on repository poisoning...
arXiv:2609.39065v1 Announce Type: cross Abstract: LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabi...
arXiv:2608. 14876v1 Announce Type: cross Abstract: Agentic coding assistants are finding widespread use, not just in new code development but in quickly ingesting and leveraging third-party code.
arXiv:2608.30686v1 Announce Type: cross Abstract: Coding agents are increasingly used for software engineering tasks, including bootstrapping projects from third-party repositories whose integrity ca...
The paper examines the aftermath of a rapid surge in AI agent skills following the viral spread of the OpenClaw AI agent in early 2026. It analyzes Git history, GitHub issues, and registry snapshots to show that the top 10% of skills dominated downloads, yet most skills lacked human review and many contained privilege‑evidencing code. Automated security scanners were inconsistent, with low sensitivity after human adjudication, highlighting the inadequacy of simple metadata or single‑scanner approaches for governing fast‑growing skill registries.
arXiv:2606. 18619v1 Announce Type: cross Abstract: The advent of agentic vulnerability detection is already becoming a watershed moment for software security.
arXiv:2609.15939v1 Announce Type: cross Abstract: Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can dete...
The paper examines the rapid growth of the OpenClaw AI agent’s public skill registry, noting a near doubling of the observable stock in 91 days and a concentration of activity in a short period. It finds that only a small fraction of skills receive significant attention—most have no stars or comments—while a large portion contains privilege‑evident code. Automated security scanners show low agreement and sensitivity, indicating that current tools are insufficient for reliable governance of fast‑expanding agent‑skill ecosystems.
arXiv:2609.36879v1 Announce Type: cross Abstract: As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent...