arXiv AI

ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution

arXiv:2608. 05790v1 Announce Type: new Abstract: General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain environments.

arXiv AI
Aug 20

LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents

LEDGER is a tracing and review system for large language model agents that constructs layered trace graphs from observed sessions. It groups raw trace records into Evidence Nodes and Workflow Nodes, anchors artifacts as evidence, and adds typed semantic edges linking claims to supporting actions, artifacts, and checks. The resulting traces reveal workflow decisions, artifact lineage, repair steps, validation coverage, and claim‑support paths for evidence‑centered audit.

By Daehong Kim, Haichao Miao, Shusen Liu
arXiv AI
Aug 19

When Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling

The paper surveys the growing use of AI agents that modify external state via the Model Context Protocol (MCP) ecosystem, noting an increase from 27% to 65% of tool use. It argues that when such agents operate on public blockchains, the blockchain execution layer’s properties—irreversibility, signing authority, continuous autonomy, and sequence-level composition—reshape the threat model, making failures irreversible. The authors organize existing MCP-security literature into an attack-surface taxonomy, present a Web3 risk-mapping matrix linking attack classes to amplified impacts and mitigations, and conclude that current defenses are inadequate, stopping fewer than 30% of attacks and less than 3% of model-level safety failures. whyItMatters":"The study highlights that AI agents acting on Web3 introduce irreversible risks that conventional software security cannot address, underscoring the need for stronger, blockchain-aware safeguards."

By Rabimba Karanjai (Larry), Yang Lu (Larry), Nour Diallo (Larry), Wujie Xiong (Larry), Lei Xu (Larry), Weidong (Larry), Shi
arXiv AI
Sep 17

Symbolic Temporal Supervision of LLM Agents Using Contracts

ContrAgent is a contract‑based framework that provides symbolic temporal supervision for large language model agents. It records an agent’s tool‑call sequence as a trace of checkable predicates and formalizes desired behaviors with assume‑guarantee contracts expressed in linear temporal logic over finite traces (LTLf). Each contract is compiled into a deterministic finite automaton that both gates actions online and evaluates recorded traces offline, enabling deterministic, reproducible verdicts and significantly lower per‑call latency compared to existing LLM‑judge and rule‑based guardrail baselines.

By Yifeng Xiao, Pierluigi Nuzzo
arXiv AI
Jul 16

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

arXiv:2607. 13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical.

By Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu