arXiv:2606. 26028v2 Announce Type: replace-cross Abstract: As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent assess whether an unknown counterpart is trustworthy?
By Xihan Xiong, Zelin Li, Wei Wei, Qin Wang, William Knottenbelt, Zhipeng Wang
As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent assess whether an unknown counterpart is trustworthy? The ERC-8004 protocol addresses this challenge with the first permissionless trust layer for AI agent economies, built around three on-chain registries for Identity, Reputation, and Validation.
CAVEAT is a new benchmark that tests computer‑use agents (CUAs) in nine online marketplace environments where platform incentives may steer agents away from user goals. The study finds that agents succeed in choosing user‑optimal products only 78.6% of the time in neutral settings, dropping to 17.3% when steering mechanisms are active. By diagnosing three failure points—priority distortion, premature narrowing of options, and early commitment—CAVEAT-Harness interventions raise user‑optimal purchasing success by 55.0%.
By Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
arXiv:2607. 13998v1 Announce Type: cross Abstract: The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms.
By Sai Srikanth Madugula, Peplluis Esteva de la Rosa, Daya Shankar
arXiv:2606. 24783v1 Announce Type: cross Abstract: Commercial NLP treats the shopping chatbot as a recommender or a conversion tool: its job is to match a user to a catalogue entry and close a sale.
By Filippos Ventirozos, Matthew Shardlow
Commercial NLP treats the shopping chatbot as a recommender or a conversion tool: its job is to match a user to a catalogue entry and close a sale. We argue that the arrival of agent-native micro-payment rails (e.
arXiv:2609.22601v1 Announce Type: cross
Abstract: The increasing reliance of autonomous AI agents on external and distributed knowledge sources introduces a fundamental challenge for decentralized in...
By Yixiang Yao, Pasha Barahimi, Srivatsan Ravi
StableEval Arena is a cost‑aware benchmark framework designed to evaluate agentic AI systems on stablecoin peg‑risk prediction. It tests LLM‑backed agents by diagnosing peg stress and forecasting deviations from the one‑dollar peg over a hidden seven‑day horizon, using leakage‑safe historical replay with exchange price‑volume data and market‑context features. The benchmark includes a 120‑case stress‑enriched validation block and a 507‑case natural‑distribution full‑arena evaluation, measuring prediction quality, calibrated‑label behavior, structured‑output reliability, latency, token consumption, and estimated inference cost across six LLM‑backed agent configurations and baselines.
By Sean Wan, Dongping Liu, Luyao Zhang
arXiv:2606. 13608v1 Announce Type: new Abstract: Agent systems are advancing quickly across domains, but their evaluation remains fragmented.
By Xiaoyuan Liu, Jianhong Tu, Yuqi Chen, Siyuan Xie, Sihan Ren, Tianneng Shi, Gal Gantar, Evan Sandoval, Donghyun Lee, Daniel Miao, Peter J. Gilbert, Nick Hynes, Mauro Staver, Warren He, David Marn, Andrew Low, Xi Zhang, Elron Bandel, Michal Shmueli-Scheuer, Siva Reddy, Alexandre Drouin, Alexandre Lacoste, Ramayya Krishnan, Elham Tabassi, Yu Su, Victor Barres, Chenguang Wang, Wenbo Guo, Dawn Song
arXiv:2605. 25815v4 Announce Type: replace Abstract: Agent-to-Agent (A2A) networks enable autonomous AI agents to collaborate by sharing reusable problem-solving instructions.
By Qiming Ye, Peixian Zhang, Yupeng He, Zifan Peng, Gareth Tyson
The paper proposes a scalable trust discovery architecture for the Internet of Agents, featuring a three‑layer hierarchical design: Agent Root for registry governance, Agent Registry for registration and metadata, and Agent Resolver for capability discovery. It introduces a registry‑suffix‑anchored composite identity scheme and a dual‑certificate, multi‑level authentication mechanism to strengthen agent identity trust. Prototype evaluation shows low latency (58 ms registration, 25 ms discovery) and high throughput (over 19,000 registrations and 29,000 discoveries per second).
By Song Zhang, Jiankang Yao, Hongtao Li, Xiaojun Zhang, Xugang Shen, Xin Li, Yanbiao Li
arXiv:2504. 16116v4 Announce Type: replace-cross Abstract: The Web3 ecosystem, underpinned by cryptographic primitives and decentralized consensus, represents a high-stakes environment where software vulnerabilities and incentive misalignments translate directly into financial loss.
By Enhao Huang, Pengyu Sun, Shuxun Wang, Zixin Lin, Alex Chen, Kaichun Hu, Joey Ouyang, Frank Li, Zhiyu Zhang, Haobo Wang, Yiming Li, Zhan Qin, James Yi, Gang Zhao, Ziang Ling, Lowes Yang