arXiv AI By Giulio Zingrillo, Hanna Foerster, Ilia Shumailov, Yiren Zhao, Robert Mullins

Securing Computer-Use Agents Against Branch Steering Attacks

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 23

ActGov: Governing LLM Agent Actions via Policy-Constrained Validation

ActGov is a runtime enforcement framework that validates each action proposed by a large language model (LLM) agent before it interacts with external tools, ensuring that actions stay within task‑scoped authorization boundaries and comply with dynamically constructed policies. It builds policies from tool specifications, benign tasks, and failure traces, verifying updates via SMT‑based counterexample checking. In evaluations on AgentDojo and AgentDyn benchmarks, ActGov consistently reduces indirect prompt‑injection attack success while maintaining task utility, outperforming existing defenses.

By Kaiyuan Zhang, Yuke Peng, Ke Jiang, Yinqian Zhang
arXiv AI
Aug 26

WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents

WebMCP-Phalanx introduces a dual‑layer runtime for browser‑integrated LLM agents that enforces trust boundaries on web‑exposed tools. The first layer uses cryptographic capability credentials to bind tools to their registering principals and propagate provenance labels, while the second layer separates semantic inspection from privileged tool use via a Quarantine Agent that validates tool metadata before a Privileged Agent can execute it. Empirical results show the approach eliminates revocation and overwrite attacks, blocks most prompt‑injection attempts, and maintains task utility comparable to a no‑attack baseline.

By Lin-Fa Lee, YI-YU Chang, Kuo-Hui Yeh