arXiv AI

terms.txt: A Consent and Compensation Protocol for Agentic Web Access

The paper introduces **terms.txt**, a new protocol that extends the traditional robots.txt format to allow website owners to specify per‑path, per‑purpose access rules for automated agents. It proposes an origin‑enforced exchange using Web Bot Auth signatures, signed intent, delegation tokens, HTTP 402 negotiation, and signed receipts to enforce, audit, and contract these rules. A lightweight implementation adds only 0.20 to 0.65 ms per request on a single vCPU.

arXiv AI
Jun 18

Towards an Agent-First Web: Redesigning the Web for AI Agents

arXiv:2606. 19116v1 Announce Type: new Abstract: The World Wide Web was built on an assumption held for three decades: the primary consumer of web content is a human being.

By Eranga Bandara, Ross Gore, Ravi Mukkamala, Asanga Gunaratna, Safdar H. Bouk, Xueping Liang, Peter Foytik, Abdul Rahman, Sachini Rajapakse, Isurunima Kularathna, Pramoda Karunarathna, Chalani Rajapakse, Ng Wee Keong, Kasun De Zoysa, Tharaka Hewa, Amin Hass, Wathsala Herath, Aruna Withanage, Nilaan Loganathan, Atmaram Yarlagadda, Sachin Shetty
arXiv AI
Sep 4

Identifying AI Web Scrapers Using Canary Tokens

The paper introduces a method to detect which web scrapers feed data to large language models (LLMs) by deploying dynamic websites that issue unique canary tokens to each scraper. By querying LLMs for information about these sites, the authors can identify when an LLM consistently outputs the unique tokens, indicating exposure to a specific scraper. Experiments on 22 production LLM systems show the technique reliably uncovers both known and undisclosed scrapers, offering a tool for third parties to monitor and control unwanted web scraping.

By Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger
arXiv AI
Jul 28

Intent-Governed Tool Authorization for AI Agents

arXiv:2606. 22916v2 Announce Type: replace Abstract: AI agents increasingly act through external tools: they read private data, construct structured payloads, submit write requests, export records, and coordinate workflows across application boundaries.

By Genliang Zhu, Chu Wang
arXiv AI
Aug 19

Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

The paper introduces Aegis, a runtime governance system for agentic AI that treats model outputs as action proposals and mediates them through a trusted decision layer before tool execution. Aegis evaluates proposals against active policy, resolves provenance server‑side, fails closed under uncertainty, and routes selected cases through a Senate‑style settlement process. In a sandbox evaluation across 6,300 rows, Aegis prevented all governed mock‑tool applications and risky side‑effect completions, preserving provenance and quorum evidence for all settled cases.

By Adam Mazzocchetti
arXiv AI
Aug 26

WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents

WebMCP-Phalanx introduces a dual‑layer runtime for browser‑integrated LLM agents that enforces trust boundaries on web‑exposed tools. The first layer uses cryptographic capability credentials to bind tools to their registering principals and propagate provenance labels, while the second layer separates semantic inspection from privileged tool use via a Quarantine Agent that validates tool metadata before a Privileged Agent can execute it. Empirical results show the approach eliminates revocation and overwrite attacks, blocks most prompt‑injection attempts, and maintains task utility comparable to a no‑attack baseline.

By Lin-Fa Lee, YI-YU Chang, Kuo-Hui Yeh