arXiv AI

Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents

arXiv:2608. 12273v1 Announce Type: cross Abstract: LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning.

arXiv AI
Sep 10

AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing

The paper introduces AgentLeak, a black‑box attack that clones the task‑solving capabilities of a strong LLM agent onto a weaker one by exploiting differences between successful and failed executions. Unlike prior skill‑stealing methods that only recover explicit skill artifacts, AgentLeak identifies and incorporates missing procedural behaviors, boosting task pass rates by over 40% and closing more than 80% of the capability gap across 20 scenarios. The study demonstrates that observable execution behavior can leak proprietary procedural knowledge, posing a confidentiality risk for LLM agents.

By Xiaoting Lyu, Yuhong Wu, Yufei Han, Shichang Liu, Liang Zhang, Bin Wang, Bin Wang, Xiaobo Ma, Wei Wang
arXiv AI
2d ago

Chaining Skills to Hijack LLM Agents

arXiv:2610.01564v1 Announce Type: cross Abstract: LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowi...

By Tian Dong, Zixuan Ma, Haodong Zhao, Huaien Zhang, Shaofeng Li, Hao Chen
arXiv AI
3d ago

Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents

The paper introduces a new skill poisoning technique for large language model agents that decouples the pretext (rationale) from the actuation (operation). By separating these two risk‑realization factors, the authors create coordinated pretext‑actuation skill pairs that allow malicious actions to remain hidden within legitimate agent behavior. An automated framework is presented to discover execution dependencies, synthesize these skill pairs, and refine them through closed‑loop feedback, achieving high attack success in both single‑session and persistent scenarios.

By Wenxin Wu, Lingyong Yan, Lei Sha, Shuaiqiang Wang, Jiashu Zhao
arXiv Computation and Language
Aug 25

SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents

The paper "SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents" investigates how agent skills—task‑specific instructions, scripts, and resources—can be exploited to create a trusted instruction channel that enables token amplification attacks. It introduces a two‑phase framework, SkillBloat, which first screens a library of attack‑type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM‑guided full‑document skill rewriting. Evaluated on a real‑world skill benchmark, SkillBloat achieves an average best amplification of 5.4184×–10.1455× across multiple coding‑agent target configurations, and an ablation study shows that the second‑stage refinement consistently improves performance over the initial screening alone.

By Yuanjin Zheng, Jingbang Chen
arXiv AI
Aug 24

ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents

ClawSentry is an open‑source, framework‑agnostic security supervision gateway designed to protect autonomous large language model (LLM) agents from progressive risks that can arise at four points in the agent control loop: skill admission, invocation‑time intent, execution‑time effect, and post‑action consequence. It introduces a multi‑tier decision engine—deterministic L1, rule‑anchored L2, and read‑only L3—alongside a First‑Use Skill Package Review (FSPR) and an Agent Harness Protocol (AHP) that applies a single policy across multiple agent runtimes without modifying their internals. Evaluation on SkillInject and the SkillsSafety benchmark shows that ClawSentry significantly reduces contextual adversarial skill risk (ASR) while maintaining high task success rates (TSR).

By Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu
arXiv AI
Sep 15

Overflip: Repetition-Induced Label Flips in Guardrail Models

Guardrail models, which screen malicious prompts in LLM services, often use lightweight Transformers with short context windows and bucketed positional encodings. The study identifies a new failure mode called Overflip, where repeating a prompt causes the guardrail’s prediction to flip from malicious to benign as the sequence length increases. Experiments on nine popular guardrails show that 5 models exhibit MAL→BEN flips on 100 prompts, with flip rates ranging from 8% to 92% and first flips occurring between 2.6k and 9.4k tokens, highlighting a gradual attention dispersion distinct from traditional attention‑dilution attacks.

By Xu He, Chih-Hsuan Lin, Hung-Mao Chen, Junjie Xiong, Yan Zhai, Kun Sun