E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.
By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
arXiv:2605. 13909v2 Announce Type: replace-cross Abstract: Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation.
By Erica Zhang, Fangzhao Zhang, Aneesh Pappu, Batu El, Jose Blanchet, Susan Athey, Jiashuo Liu, James Zou
arXiv:2608. 08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work.
By Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing
arXiv:2608. 10475v1 Announce Type: new Abstract: The emergence of language-based AI agents promises to transform the scope of machine economic activity.
By Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann, Tan Zhi-Xuan
arXiv:2608. 14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation.
By Wael Albayaydh, Rui Zhao
The paper investigates how marketplace guardrails affect welfare in language‑model agent simulations of hotel transactions. It finds that initial reports of large welfare gains disappear when controlling for offer schemas and buyer choice, and that guardrails mainly redistribute rather than increase welfare unless sellers are explicitly forced to produce inefficient bundles. The authors propose a construct‑validity framework to flag invalid or inconclusive policy claims before they are reported.
By Peiying Zhu, Sidi Chang