E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.
By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
arXiv:2605. 13909v2 Announce Type: replace-cross Abstract: Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation.
By Erica Zhang, Fangzhao Zhang, Aneesh Pappu, Batu El, Jose Blanchet, Susan Athey, Jiashuo Liu, James Zou
arXiv:2608. 08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work.
By Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing
arXiv:2608. 10475v1 Announce Type: new Abstract: The emergence of language-based AI agents promises to transform the scope of machine economic activity.
By Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann, Tan Zhi-Xuan
arXiv:2608. 14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation.
By Wael Albayaydh, Rui Zhao
The paper investigates how marketplace guardrails affect welfare in language‑model agent simulations of hotel transactions. It finds that initial reports of large welfare gains disappear when controlling for offer schemas and buyer choice, and that guardrails mainly redistribute rather than increase welfare unless sellers are explicitly forced to produce inefficient bundles. The authors propose a construct‑validity framework to flag invalid or inconclusive policy claims before they are reported.
By Peiying Zhu, Sidi Chang
arXiv:2607. 05863v1 Announce Type: new Abstract: Negotiation is a fundamental strategic interaction in management science, characterized by agents attempting to reach agreements while protecting private information, such as reservation costs and hidden valuations.
By Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao, Xin Chen, David Simchi-Levi
arXiv:2606. 01456v1 Announce Type: new Abstract: Large language models are increasingly deployed as advisors whose objective is not aligned with the user's: recommenders optimize for engagement, sales assistants for purchases, negotiation agents for concessions.
By Hamidreza Hasani Balyani, Seyed Pouyan Mousavi Davoudi, Alireza Amiri-Margavi, Amin Gholami Davodi, Arshia Gharagozlou
arXiv:2606. 03034v1 Announce Type: cross Abstract: Large language model (LLM) agents have begun to delegate work to one another.
By Gaurav Naresh Mittal
arXiv:2607. 02814v1 Announce Type: cross Abstract: Personal agents will increasingly negotiate on behalf of users: splitting costs with other personal agents, appealing platform decisions, escalating support disputes, requesting refunds, changing subscriptions, and negotiating deadlines or reimbursements.
By Dylan Zongmin Liu
arXiv:2607. 28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria.
By Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo
PriceBench is a diagnostic benchmark that extracts price, quality, and brand preferences from large language models (LLMs) by analyzing their hotel booking choices. Using a logit choice model, the study evaluated 28 LLMs from eight providers across 3,600 booking tasks involving 179 New York City hotels. Results show that more capable LLMs exhibit stronger, more consistent preferences, while weaker models either lock onto a single position or show near-indifference, with significant variation in price sensitivity and price/quality trade-offs across providers.
By Pavel Kireyev