arXiv AI

Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering

arXiv:2604. 11430v2 Announce Type: replace-cross Abstract: AI agents that pay for resources via the x402 protocol embed payment metadata - resource URLs, descriptions, and reason strings - in every HTTP payment request.

arXiv AI
Aug 19

PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance

The paper introduces PACE (Policy‑Attested Contract Execution), a framework that sits between large‑language‑model (LLM) based autonomous AI agents and on‑chain DeFi operations. PACE defines typed transaction intents, a deterministic policy verifier, and signed Policy Decision Records (PDRs) that cryptographically bind an approved intent, policy, and simulation report to the exact on‑chain execution bytes, providing replay and expiration protection. In evaluations across 40 tasks and six baselines, PACE achieves zero unsafe executions and zero false positives, outperforming unguarded agents by a large margin.

By Rabimba Karanjai (Larry), Yang Lu (Larry), Richard Williamson (Larry), Hemanth Hm (Larry), Prakhar Mehrotra (Larry), Lei Xu (Larry), Weidong (Larry), Shi
arXiv AI
Sep 2

A Formal Analysis of Agent Payment Protocols

The paper presents a formal analysis of four agent payment protocols—x402, MPP, ACP, and AP2—using the Tamarin prover. By modeling each protocol’s roles, state, and trust assumptions, the authors verify 86 cases, reproducing 46 known results and uncovering 40 new formal-consistency findings. They further validate ten findings through implementation proofs of concept, SDK/schema witnesses, and executable traces, highlighting the importance of consistent delegated authorization across all protocol stages.

By Ke Jiang, Mohan Yu, Yuan Chang, Mohit Kumar Jangid, Jianyu Niu, Cong Wang, Yinqian Zhang
arXiv AI
Sep 15

AcquireBound: Runtime Authorization for Resources Acquired by AI Agents

AcquireBound is a runtime authorization framework that ensures AI agents can safely acquire and activate resources such as compute, credentials, and services. It quarantines acquired outputs, resolves their capabilities through authenticated evidence, and activates them only after verifying a manifest, provenance, and relational constraints. The system demonstrates strong safety properties, passing extensive benign and unsafe trace tests across multiple resource classes.

By Genliang Zhu
arXiv AI
Jun 12

The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements

arXiv:2606. 12797v1 Announce Type: new Abstract: Agentic large language model systems that autonomously invoke tools, maintain persistent memory, and execute multi-step plans are increasingly deployed in public-facing domains, including government services, healthcare triage, and financial advising.

By Md Jafrin Hossain, Mohammad Arif Hossain, Weiqi Liu, Nirwan Ansari
arXiv AI
Sep 30

Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money

The paper introduces the Agentic Commerce Bench (ACB), a benchmark for measuring fraud in AI agents that autonomously spend money. It presents a taxonomy of agentic commerce fraud, a dataset of twenty fraud classes derived from real production data, and an open‑source detector stack called gordonguard for auditing and replaying hostile counterparties. The study shows that current reasoning layers and security scanners perform poorly on many classes, highlighting the need for better detection mechanisms.

By Ankit Srivastava, Debjyoti Paul
arXiv AI
Aug 26

Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)

The paper presents a systematic security analysis of Google’s Agent Payments Protocol (AP2) version 2.0, focusing on its roles, transaction lifecycle, and deployment architectures. It identifies 48 threats across five attack families, scores them with the AIVSS, and demonstrates eight high‑risk threats with proof‑of‑concept attacks and mitigations. The study also introduces a deployment‑aware scanner to map threats to various checks, showing that signed mandates alone cannot guarantee user intent when pre‑authorization context is manipulated.

By Avital Aviv, Parth A. Gandh, Ron Bitton, Asaf Shabtai
arXiv AI
Sep 17

BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

BENCHCOMPASS is a new payment‑domain benchmark that transforms typed evidence packs into scenario‑grounded tasks, applies LLM‑based quality checks, generates attack variants, and reserves final item admission for domain experts. It includes an expert‑reviewed Pro benchmark covering payment knowledge, context‑grounded scenario reasoning, and attacked open robustness, plus a lower‑assurance Normal pool. Across 16 model variants, BENCHCOMPASS reveals distinct failure modes—missing payment knowledge, incomplete reasoning, and failure to reject invalid workflows—while the best model scores 89.6% on open context‑grounded reasoning and 81.7% under attacked inputs. "whyItMatters":"The benchmark provides a structured way to isolate and evaluate specific weaknesses in LLMs for payment operations, a critical financial infrastructure where rules change rapidly and decisions depend on complex contextual factors."

By Sijie Dong, Wei Ren, Xuanwei Hu, Jiawei Luo, Zifan Wang, Xiaoyun Feng, Hui Cai, Lyuxin Xue, Peng Lu, Jianshe Li, Xin Zhang, Wei Wu
arXiv Machine Learning
Aug 24

If It Walks Like an Arbitrage: Protocol-Agnostic Detection with Decidable Structural Equivalence

The paper presents a protocol‑agnostic method for detecting arbitrage in Ethereum by converting transaction traces into a canonical abstract syntax tree using a convergent rewriting system of 15 rules. This canonical form enables decidable structural equivalence of fund flows, allowing the authors to identify arbitrage cycles without relying on protocol‑specific patterns. Evaluated on 220,000 Ethereum blocks, the system confirmed 469,801 arbitrage opportunities, matching 83.5% of a production MEV platform and covering 81% of a GNN classifier, while producing no false positives in a manual sample.

By Adam Khayam, Hamid Kolli, Mohamed Iguernalala, \c{C}agdas Bozman