arXiv AI

EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making

arXiv AI
Sep 7

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

ERPBench is a benchmark that evaluates large language model agents in enterprise decision-making through a six‑round ERP simulation covering pricing, production, procurement, inventory, finance, and market competition. It tests the same 100 problems in two market ecologies—Solo, where agents compete against rule‑based opponents, and Arena, where six agents compete together—producing 1,200 model trajectories across 7,200 decision rounds. Results show that model performance varies by ecology, with DeepSeek best in Solo and Gemini best in Arena, and only 21 of 100 problems yield the same top performer across both settings.

By Xinran Zhang, Pengrui Lu, Lyumanshan Ye, Pengfei Liu
arXiv AI
Jul 21

Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning

arXiv:2607. 17331v1 Announce Type: new Abstract: Enterprise Resource Planning (ERP) systems record transactions reliably but still delegate almost all operational decision-making to human specialists, because classical rule-based automation cannot reason about exceptions and monolithic AI assistants degrade when asked to coordinate across functional boundaries.

By Zhihao Liu, Tianyu Wang, Xi Vincent Wang, Lihui Wang
arXiv AI
Jul 28

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

arXiv:2607. 23124v1 Announce Type: new Abstract: Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings.

By Hao Jiang, Gangtao Xin, Yingdi Huang, Guojie Zhu, Jiangshan Zhang, Xinyuan Lin, Yunkun Xu, Chengyu Shen, Wenlong Fei, Jiawei Li, Yujie Fu, Sichen Kang, Tingyu Xie, Yedi Hu, Jingren Zhang, Hongcheng Gao, Jianshu Zeng, Chong Chen, Chang Guo, Chao Feng, Feng Wang, Fulin Lin, Jinchao Ma, Lang Mei, Li Huang, Liyan Liu, Qing He, Shuting Tao, Siyu Mo, Xiangnan Chen, Xiaohan Yu, Xiaoyang Li, Yanheng Hou, Yanyu Wu, Zhihan Yang, Wentao Zhang, Yang Gao, Zhao Cao
arXiv AI
Jun 2

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

arXiv:2606. 02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks.

By Shannon Serrao, Soumitra Chatterjee, Dorina Strori, Abhishek Sharma, Nathan Miller