arXiv AI

LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces

Hugging Face Trending Papers
Jun 24

Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem

As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent assess whether an unknown counterpart is trustworthy? The ERC-8004 protocol addresses this challenge with the first permissionless trust layer for AI agent economies, built around three on-chain registries for Identity, Reputation, and Validation.

arXiv AI
3d ago

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

CAVEAT is a new benchmark that tests computer‑use agents (CUAs) in nine online marketplace environments where platform incentives may steer agents away from user goals. The study finds that agents succeed in choosing user‑optimal products only 78.6% of the time in neutral settings, dropping to 17.3% when steering mechanisms are active. By diagnosing three failure points—priority distortion, premature narrowing of options, and early commitment—CAVEAT-Harness interventions raise user‑optimal purchasing success by 55.0%.

By Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
arXiv AI
Sep 17

StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction

StableEval Arena is a cost‑aware benchmark framework designed to evaluate agentic AI systems on stablecoin peg‑risk prediction. It tests LLM‑backed agents by diagnosing peg stress and forecasting deviations from the one‑dollar peg over a hidden seven‑day horizon, using leakage‑safe historical replay with exchange price‑volume data and market‑context features. The benchmark includes a 120‑case stress‑enriched validation block and a 507‑case natural‑distribution full‑arena evaluation, measuring prediction quality, calibrated‑label behavior, structured‑output reliability, latency, token consumption, and estimated inference cost across six LLM‑backed agent configurations and baselines.

By Sean Wan, Dongping Liu, Luyao Zhang
arXiv AI
Jun 12

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

arXiv:2606. 13608v1 Announce Type: new Abstract: Agent systems are advancing quickly across domains, but their evaluation remains fragmented.

By Xiaoyuan Liu, Jianhong Tu, Yuqi Chen, Siyuan Xie, Sihan Ren, Tianneng Shi, Gal Gantar, Evan Sandoval, Donghyun Lee, Daniel Miao, Peter J. Gilbert, Nick Hynes, Mauro Staver, Warren He, David Marn, Andrew Low, Xi Zhang, Elron Bandel, Michal Shmueli-Scheuer, Siva Reddy, Alexandre Drouin, Alexandre Lacoste, Ramayya Krishnan, Elham Tabassi, Yu Su, Victor Barres, Chenguang Wang, Wenbo Guo, Dawn Song
arXiv AI
Sep 18

A Scalable Trust Discovery Architecture for the Internet of Agents

The paper proposes a scalable trust discovery architecture for the Internet of Agents, featuring a three‑layer hierarchical design: Agent Root for registry governance, Agent Registry for registration and metadata, and Agent Resolver for capability discovery. It introduces a registry‑suffix‑anchored composite identity scheme and a dual‑certificate, multi‑level authentication mechanism to strengthen agent identity trust. Prototype evaluation shows low latency (58 ms registration, 25 ms discovery) and high throughput (over 19,000 registrations and 29,000 discoveries per second).

By Song Zhang, Jiankang Yao, Hongtao Li, Xiaojun Zhang, Xugang Shen, Xin Li, Yanbiao Li
arXiv AI
Jun 29

DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain

arXiv:2504. 16116v4 Announce Type: replace-cross Abstract: The Web3 ecosystem, underpinned by cryptographic primitives and decentralized consensus, represents a high-stakes environment where software vulnerabilities and incentive misalignments translate directly into financial loss.

By Enhao Huang, Pengyu Sun, Shuxun Wang, Zixin Lin, Alex Chen, Kaichun Hu, Joey Ouyang, Frank Li, Zhiyu Zhang, Haobo Wang, Yiming Li, Zhan Qin, James Yi, Gang Zhao, Ziang Ling, Lowes Yang