arXiv AI By Min Yang, Yichen Pan, Jinghua Piao, Dandan Song, Yongshun Gong, Yong Li

EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 7

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

ERPBench is a benchmark that evaluates large language model agents in enterprise decision-making through a six‑round ERP simulation covering pricing, production, procurement, inventory, finance, and market competition. It tests the same 100 problems in two market ecologies—Solo, where agents compete against rule‑based opponents, and Arena, where six agents compete together—producing 1,200 model trajectories across 7,200 decision rounds. Results show that model performance varies by ecology, with DeepSeek best in Solo and Gemini best in Arena, and only 21 of 100 problems yield the same top performer across both settings.

By Xinran Zhang, Pengrui Lu, Lyumanshan Ye, Pengfei Liu
arXiv AI
Jul 21

Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning

arXiv:2607. 17331v1 Announce Type: new Abstract: Enterprise Resource Planning (ERP) systems record transactions reliably but still delegate almost all operational decision-making to human specialists, because classical rule-based automation cannot reason about exceptions and monolithic AI assistants degrade when asked to coordinate across functional boundaries.

By Zhihao Liu, Tianyu Wang, Xi Vincent Wang, Lihui Wang