arXiv:2609.37658v1 Announce Type: new
Abstract: LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-te...
By Min Yang, Yichen Pan, Jinghua Piao, Dandan Song, Yongshun Gong, Yong Li
arXiv:2608. 08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work.
By Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing
ERPBench is a benchmark that evaluates large language model agents in enterprise decision-making through a six‑round ERP simulation covering pricing, production, procurement, inventory, finance, and market competition. It tests the same 100 problems in two market ecologies—Solo, where agents compete against rule‑based opponents, and Arena, where six agents compete together—producing 1,200 model trajectories across 7,200 decision rounds. Results show that model performance varies by ecology, with DeepSeek best in Solo and Gemini best in Arena, and only 21 of 100 problems yield the same top performer across both settings.
By Xinran Zhang, Pengrui Lu, Lyumanshan Ye, Pengfei Liu
arXiv:2605. 22664v2 Announce Type: replace Abstract: LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions.
By Thomson Yen, Julian Poeltl, Harshith Srinivas Gear, Yilin Meng, Joshua Fan, Adam Shen, Yili Liu, Ali Bauyrzhan, Siri Du, Haoyang Liu, Daniel Guetta, Hongseok Namkoong
E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.
By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
arXiv:2606. 18543v1 Announce Type: new Abstract: Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service.
By Haozhe Chen, Karthik Narasimhan, Zhuang Liu