Query-Only Backdoor Attacks on Self-Evolving Skills via Trajectory Poisoning
arXiv:2608. 08303v1 Announce Type: new Abstract: Agentic skills improve large language model (LLM) agents by encoding reusable procedures for complex tasks.
arXiv:2608. 09577v1 Announce Type: new Abstract: Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it.
arXiv:2608. 08303v1 Announce Type: new Abstract: Agentic skills improve large language model (LLM) agents by encoding reusable procedures for complex tasks.
arXiv:2606. 07943v1 Announce Type: cross Abstract: Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks.
arXiv:2608. 09732v1 Announce Type: cross Abstract: Agent skills are emerging as an important attack surface in LLM-based agent systems.
The paper introduces Quarantined Expert Shutdown (QES), a new backdoor containment strategy for large language models. QES allows backdoor learning to occur during training but routes it into a designated, quarantined expert that can be disabled at deployment. The method achieves significant reductions in attack success rates while largely preserving model utility.
arXiv:2602. 14211v3 Announce Type: replace-cross Abstract: Agent skills extend LLM agents with task-specific instructions, executable scripts, and auxiliary resources, improving reusability but creating a new supply-chain attack surface.
arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.
The paper proposes universal, tool‑based defenses for large language model agents that use external tools, addressing four types of adversarial attacks: direct and indirect prompt injection, memory poisoning, and backdoor attacks. Two main defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to remove suspicious tools, and Normal Tool Recalling, which restores the agent’s original toolset before planning. The authors also add prompt‑based defenses such as Chain‑of‑Thought prompting and self‑reflection, and demonstrate that these methods dramatically lower attack success rates—often to 0%—across multiple open‑source and proprietary LLMs while maintaining or improving task performance.
arXiv:2609.36570v1 Announce Type: cross Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that...
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.
arXiv:2609.01487v1 Announce Type: cross Abstract: Skill-augmented agents load reusable skills as persistent runtime context, improving task performance but also giving malicious skills a durable chan...
arXiv:2608.30041v1 Announce Type: cross Abstract: Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later pri...
The paper introduces skill cascading attacks, where a malicious goal is spread across multiple seemingly benign skills, causing harmful outcomes when combined. It presents SkillCascade, an automated red‑teaming framework, and releases SkillCascade‑Bench, a benchmark of 213 validated cascading test cases across various agent systems and domains. Experiments show that these cascaded interactions reliably induce harmful behaviors while evading existing per‑skill scanners and runtime monitors, revealing a gap between component‑level integrity and system‑level safety.