arXiv AI By Jinqian Zhang (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences), Haojun Xia (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences), Shujiang Wu (Beihang University), Jingkun Yue (State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China), Xia Zhang (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences), Zhangpei Cheng (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences), Bibo Tu (Institute of Information Engineering, Chinese Academy of Sciences, School of Cyber Security, University of Chinese Academy of Sciences)

Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents

Read the original on arXiv AI →

The paper investigates how multi‑step tool‑calling large language model agents can unintentionally create persistent billable state, allowing malicious or compromised tools to generate repeated charges without user credentials. It formalizes the persistent billable‑state boundary, identifies six denial‑of‑wallet attack vectors, and evaluates them with the DOW‑BENCH harness across six model families. The study shows significant cost amplification, demonstrates effective mitigation via deterministic history transformation and host‑side invariants, and highlights the scarcity of existing safeguards in real‑world repositories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 23

ActGov: Governing LLM Agent Actions via Policy-Constrained Validation

ActGov is a runtime enforcement framework that validates each action proposed by a large language model (LLM) agent before it interacts with external tools, ensuring that actions stay within task‑scoped authorization boundaries and comply with dynamically constructed policies. It builds policies from tool specifications, benign tasks, and failure traces, verifying updates via SMT‑based counterexample checking. In evaluations on AgentDojo and AgentDyn benchmarks, ActGov consistently reduces indirect prompt‑injection attack success while maintaining task utility, outperforming existing defenses.

By Kaiyuan Zhang, Yuke Peng, Ke Jiang, Yinqian Zhang