EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
ERPBench is a benchmark that evaluates large language model agents in enterprise decision-making through a six‑round ERP simulation covering pricing, production, procurement, inventory, finance, and market competition. It tests the same 100 problems in two market ecologies—Solo, where agents compete against rule‑based opponents, and Arena, where six agents compete together—producing 1,200 model trajectories across 7,200 decision rounds. Results show that model performance varies by ecology, with DeepSeek best in Solo and Gemini best in Arena, and only 21 of 100 problems yield the same top performer across both settings.
arXiv:2609.13561v1 Announce Type: new Abstract: Efficient utilization of supply chain analytics for decision making remains a significant challenge for planners, as critical tasks such as database qu...
arXiv:2608. 08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work.
arXiv:2607. 17331v1 Announce Type: new Abstract: Enterprise Resource Planning (ERP) systems record transactions reliably but still delegate almost all operational decision-making to human specialists, because classical rule-based automation cannot reason about exceptions and monolithic AI assistants degrade when asked to coordinate across functional boundaries.
arXiv:2607. 19409v1 Announce Type: new Abstract: Recent advances in large language models have accelerated deployment of agentic systems in operational finance.
arXiv:2509. 26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions.