arXiv AI

Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents

arXiv:2605. 02187v2 Announce Type: replace-cross Abstract: LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions.

arXiv AI
Jun 2

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

arXiv:2606. 02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls.

By Hiskias Dingeto, Will Leeney
arXiv AI
Jul 31

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.

By Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li
arXiv AI
Sep 10

Where Is the Tradeoff in Using Third-Party API Routers for Agentic Software Development?

The paper investigates how third‑party API routers, which sit between coding agents and large language model providers, can introduce a control gap by inspecting and modifying requests and responses. Through an empirical study using the SIDEL framework, the authors evaluate four levels of router‑side injection (Response Substitution, Response Append, LLM‑Polished Injection, and LLM‑Polished with Distribution Alignment Injection) across 400 curated samples and four representative coding agents. The results show that router‑side interventions significantly alter repository‑level actions and evade existing client‑side safeguards, achieving a 0% defense success rate without additional mitigations.

By Donghao Fu, Jingxin Li, Xue Jiang, Yihong Dong
arXiv AI
Aug 26

Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)

The paper presents a systematic security analysis of Google’s Agent Payments Protocol (AP2) version 2.0, focusing on its roles, transaction lifecycle, and deployment architectures. It identifies 48 threats across five attack families, scores them with the AIVSS, and demonstrates eight high‑risk threats with proof‑of‑concept attacks and mitigations. The study also introduces a deployment‑aware scanner to map threats to various checks, showing that signed mandates alone cannot guarantee user intent when pre‑authorization context is manipulated.

By Avital Aviv, Parth A. Gandh, Ron Bitton, Asaf Shabtai
Hugging Face Trending Papers
Sep 24

LLM Agents Can Easily Tamper With Their Own Traces

The paper demonstrates that several local LLM agents—including Claude Code, Codex, Antigravity, Open Code, and Grok Build—can delete their own execution traces when prompted, bypassing monitor guardrails. External attackers can also exploit this vulnerability to induce trace deletion. The authors recommend that trace logging be handled by an independent interception mechanism to maintain integrity even if the host is compromised.