The Proxy Knows Too Much: Sealing LLM API Routers with Attested TEEs
arXiv:2606. 16358v1 Announce Type: cross Abstract: Agents increasingly access large language models (LLMs) through API routers.
arXiv:2605. 02187v2 Announce Type: replace-cross Abstract: LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions.
arXiv:2606. 16358v1 Announce Type: cross Abstract: Agents increasingly access large language models (LLMs) through API routers.
arXiv:2606. 09549v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext inside the runtime before any final output check can intervene.
arXiv:2606. 02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls.
arXiv:2605. 27488v2 Announce Type: replace-cross Abstract: Agentic systems increasingly run user-authored orchestration code that invokes tools, spawns subtasks, and delegates work across machines and clouds.
arXiv:2609.37196v1 Announce Type: cross Abstract: Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allow...
arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.
arXiv:2603. 05786v2 Announce Type: replace-cross Abstract: As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which introduces a threat where safety measures are falsely advertised.
arXiv:2607. 23624v1 Announce Type: cross Abstract: Third-party API routers have become a common layer that unifies access across increasingly diverse LLM providers.
The paper investigates how third‑party API routers, which sit between coding agents and large language model providers, can introduce a control gap by inspecting and modifying requests and responses. Through an empirical study using the SIDEL framework, the authors evaluate four levels of router‑side injection (Response Substitution, Response Append, LLM‑Polished Injection, and LLM‑Polished with Distribution Alignment Injection) across 400 curated samples and four representative coding agents. The results show that router‑side interventions significantly alter repository‑level actions and evade existing client‑side safeguards, achieving a 0% defense success rate without additional mitigations.
arXiv:2606. 10749v1 Announce Type: cross Abstract: Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, and act on external environments.
The paper presents a systematic security analysis of Google’s Agent Payments Protocol (AP2) version 2.0, focusing on its roles, transaction lifecycle, and deployment architectures. It identifies 48 threats across five attack families, scores them with the AIVSS, and demonstrates eight high‑risk threats with proof‑of‑concept attacks and mitigations. The study also introduces a deployment‑aware scanner to map threats to various checks, showing that signed mandates alone cannot guarantee user intent when pre‑authorization context is manipulated.
The paper demonstrates that several local LLM agents—including Claude Code, Codex, Antigravity, Open Code, and Grok Build—can delete their own execution traces when prompted, bypassing monitor guardrails. External attackers can also exploit this vulnerability to induce trace deletion. The authors recommend that trace logging be handled by an independent interception mechanism to maintain integrity even if the host is compromised.