arXiv AI

Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents

arXiv AI
Jul 7

When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents

arXiv:2607. 05189v1 Announce Type: cross Abstract: Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and proactive background execution.

By Yechao Zhang, Shiqian Zhao, Jiawen Zhang, Jie Zhang, Gelei Deng, Xiaogeng Liu, Chaowei Xiao, Tianwei Zhang
arXiv AI
Jun 12

A Survey on Long-Term Memory Security in LLM Agents: Attacks, Defenses, and Governance Across the Memory Lifecycle

arXiv:2604. 16548v2 Announce Type: replace-cross Abstract: The emergence of writable, cross-session persistent memory in LLM agents introduces a qualitatively different threat landscape from conventional input-centric security concerns, characterized by three properties: persistence, statefulness, and propagation.

By Zehao Lin, Xixuan Hao, Renyu Fu, Shaobo Cui, Kai Chen, Chunyu Li, Zhiyu Li, Feiyu Xiong
arXiv AI
Sep 25

Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents

The paper investigates how multi‑step tool‑calling large language model agents can unintentionally create persistent billable state, allowing malicious or compromised tools to generate repeated charges without user credentials. It formalizes the persistent billable‑state boundary, identifies six denial‑of‑wallet attack vectors, and evaluates them with the DOW‑BENCH harness across six model families. The study shows significant cost amplification, demonstrates effective mitigation via deterministic history transformation and host‑side invariants, and highlights the scarcity of existing safeguards in real‑world repositories.

By Jinqian Zhang (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences), Haojun Xia (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences), Shujiang Wu (Beihang University), Jingkun Yue (State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China), Xia Zhang (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences), Zhangpei Cheng (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences), Bibo Tu (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences)
arXiv AI
Jun 2

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

arXiv:2606. 02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls.

By Hiskias Dingeto, Will Leeney