Daydreaming is an execution‑only attack that steals multi‑file agent skills by interacting with a black‑box task service. By adaptively crafting tasks and analyzing the returned results, the attacker reconstructs the hidden skill without ever requesting or revealing it. In experiments on seven skills and four victim models, Daydreaming recovers 86.8% of the original capability using only 32 victim calls on average, outperforming prior methods and demonstrating that hiding skill files and filtering direct disclosure are insufficient defenses.
By Yu-Lin Tsai, Yu-An Lu, Ci-Yang Tsai, Muxi Lyu, Raluca Ada Popa, Chia-Mu Yu
The paper introduces AgentLeak, a black‑box attack that clones the task‑solving capabilities of a strong LLM agent onto a weaker one by exploiting differences between successful and failed executions. Unlike prior skill‑stealing methods that only recover explicit skill artifacts, AgentLeak identifies and incorporates missing procedural behaviors, boosting task pass rates by over 40% and closing more than 80% of the capability gap across 20 scenarios. The study demonstrates that observable execution behavior can leak proprietary procedural knowledge, posing a confidentiality risk for LLM agents.
By Xiaoting Lyu, Yuhong Wu, Yufei Han, Shichang Liu, Liang Zhang, Bin Wang, Bin Wang, Xiaobo Ma, Wei Wang
arXiv:2609.36879v1 Announce Type: cross
Abstract: As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent...
By Haoran Ou, Gelei Deng, Xuanye Zhang, Wenbo Guo, Tianwei Zhang, Kwok-Yan Lam
The paper introduces a new skill poisoning technique for large language model agents that decouples the pretext (rationale) from the actuation (operation). By separating these two risk‑realization factors, the authors create coordinated pretext‑actuation skill pairs that allow malicious actions to remain hidden within legitimate agent behavior. An automated framework is presented to discover execution dependencies, synthesize these skill pairs, and refine them through closed‑loop feedback, achieving high attack success in both single‑session and persistent scenarios.
By Wenxin Wu, Lingyong Yan, Lei Sha, Shuaiqiang Wang, Jiashu Zhao
The paper introduces skill cascading attacks, where a malicious goal is spread across multiple seemingly benign skills, causing harmful outcomes when combined. It presents SkillCascade, an automated red‑teaming framework, and releases SkillCascade‑Bench, a benchmark of 213 validated cascading test cases across various agent systems and domains. Experiments show that these cascaded interactions reliably induce harmful behaviors while evading existing per‑skill scanners and runtime monitors, revealing a gap between component‑level integrity and system‑level safety.
By Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu
arXiv:2609.39450v1 Announce Type: cross
Abstract: LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. Howe...
By Jihun Han, Yejin Jang, Byung Il Kwak, Mee Lan Han
arXiv:2609.39065v1 Announce Type: cross
Abstract: LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabi...
By Yan Wang, Zhihao Zhang, Ke Chen, Kai Chen, Yaqin Zhang, Duohe Ma, Jun Dai, Xiaoyan Sun
arXiv:2602. 06547v3 Announce Type: replace-cross Abstract: LLM-based coding agents increasingly rely on third-party extensions called skills, which bundle natural language instructions and helper scripts that execute with full user privileges.
By Yi Liu, Zhihao Chen, Yanjun Zhang, Gelei Deng, Yuekang Li, Jianting Ning, Leo Yu Zhang
arXiv:2602. 06547v4 Announce Type: replace-cross Abstract: LLM-based coding agents increasingly rely on third-party extensions called skills, which bundle natural language instructions and helper scripts that execute with full user privileges.
By Yi Liu, Zhihao Chen, Yanjun Zhang, Gelei Deng, Yuekang Li, Jianting Ning, Leo Yu Zhang
The paper "SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents" investigates how agent skills—task‑specific instructions, scripts, and resources—can be exploited to create a trusted instruction channel that enables token amplification attacks. It introduces a two‑phase framework, SkillBloat, which first screens a library of attack‑type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM‑guided full‑document skill rewriting. Evaluated on a real‑world skill benchmark, SkillBloat achieves an average best amplification of 5.4184×–10.1455× across multiple coding‑agent target configurations, and an ablation study shows that the second‑stage refinement consistently improves performance over the initial screening alone.
By Yuanjin Zheng, Jingbang Chen
arXiv:2602. 14211v3 Announce Type: replace-cross Abstract: Agent skills extend LLM agents with task-specific instructions, executable scripts, and auxiliary resources, improving reusability but creating a new supply-chain attack surface.
By Xiaojun Jia, Jie Liao, Simeng Qin, Jindong Gu, Wenqi Ren, Xiaochun Cao, Yang Liu, Philip Torr
arXiv:2608. 08264v1 Announce Type: new Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools.
By Zhengyang Shan, Xu Qian, Jiayun Xin, Kun Li, Yue Zhang, Minghui Xu