Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
arXiv:2607. 24343v1 Announce Type: cross Abstract: Language-model agents act through structured tool calls whose arguments carry different risks.
Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body but should not determine a recipient, account, command, or credential.
arXiv:2607. 24343v1 Announce Type: cross Abstract: Language-model agents act through structured tool calls whose arguments carry different risks.
The paper introduces skilder, a framework that organizes LLM agent capabilities into role‑scoped bundles of skills, tools, and instructions, with explicit limits. Agents start with a minimal role catalog, discover the roles needed for a task, and receive the associated tools only through a single MCP server, ensuring deterministic enforcement of scope. Experiments on 13 tasks with six models show that skilder’s authorization layer prevents unauthorized tool calls and parameter violations while maintaining flexibility through dynamic cross‑role capability acquisition.
ClawSentry is an open‑source, framework‑agnostic security supervision gateway designed to protect autonomous large language model (LLM) agents from progressive risks that can arise at four points in the agent control loop: skill admission, invocation‑time intent, execution‑time effect, and post‑action consequence. It introduces a multi‑tier decision engine—deterministic L1, rule‑anchored L2, and read‑only L3—alongside a First‑Use Skill Package Review (FSPR) and an Agent Harness Protocol (AHP) that applies a single policy across multiple agent runtimes without modifying their internals. Evaluation on SkillInject and the SkillsSafety benchmark shows that ClawSentry significantly reduces contextual adversarial skill risk (ASR) while maintaining high task success rates (TSR).
arXiv:2608.30041v1 Announce Type: cross Abstract: Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later pri...
arXiv:2606. 00448v1 Announce Type: cross Abstract: LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set.
arXiv:2605.12015v3 Announce Type: replace-cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files...
arXiv:2606. 18356v1 Announce Type: cross Abstract: Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persistent memory, send messages, modify databases, or trigger harmful code and tool effects.
arXiv:2606. 15899v1 Announce Type: cross Abstract: Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that extend agent capabilities - remains largely unvetted.
HarnessRisk is a lifecycle-oriented benchmark for evaluating safety in agent harnesses that manage tools, extensions, state, permissions, and external actions. It defines six operational phases—Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery—and includes 128 sandboxed cases pairing benign user objectives with adversarial instructions. Across three harnesses, six language models, and 14 configurations, attack success rates vary from 12.6% to 80.9%, with the most vulnerable phase being Harness Configuration. "whyItMatters":"The benchmark demonstrates that safety failures can arise in multiple harness responsibilities and that even explicit risk detection does not guarantee safe action, underscoring the need for comprehensive evaluation across model and harness configurations."
arXiv:2609.14780v1 Announce Type: cross Abstract: Multi-tenant tools commonly accept a tenant identifier and validate it against the caller's entitlement. For a large language model (LLM) agent, that...
LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety risks that only emerge during actual execution.
The paper introduces PACE, a Provenance-Aware Capability Enforcement system designed to secure tool-using large language model agents by mediating every tool call before execution. PACE employs path confinement to limit influence paths and verifies effects against authenticated authority, distinguishing certified execution contracts from evaluated configurations. Experiments on eight agent‑security benchmarks show that the evaluated configuration reduces attack success in most cases while maintaining near‑native utility.