arXiv AI

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

arXiv AI
Sep 11

The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

The Era by Eon Benchmark is a new dataset for evaluating large language model agents that interact with enterprise tools. It constructs a complete fictional company with product simulators, internal databases, and benchmark questions, all generated from a shared entity graph to ensure consistency. Exact answer keys are computed from the generated records, allowing precise grading and validation of realism and adversarial robustness across 23 simulated companies.

By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary
arXiv AI
Sep 17

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

ERPBench introduces a new evaluation paradigm for computer-use agents that operate via screenshots and simulated actions, focusing on enterprise software such as ERP systems. The benchmark tests agents on a live, reproducible ERP platform and scores tasks against ground-truth database values, highlighting challenges like dense interfaces, multi-step interactions, and persistent record errors. Experiments with six agents show that strong general GUI performance does not translate to reliable enterprise outcomes, with many agents frequently saving incorrect data.

By Kratika Bhagtani, Kusha Sridhar, Maziyar Baran Pouyan, Yuying Zhao, Eugene Siow
arXiv AI
Aug 28

GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions

The paper introduces GROUND, a framework that limits large language model (LLM) analytics to a governed semantic layer for enterprise data warehouses. GROUND supplies approved metrics, dimensions, join paths, filters, and security rules, then validates generated SQL against these constraints before execution, retrying or abstaining on violations. In benchmarks, GROUND eliminates hallucinations across all evaluated categories and prevents row‑level security breaches, outperforming schema‑only, schema‑RAG, and semantic‑only approaches.

By Aravind Sasidharan Pillai
arXiv AI
Sep 10

DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents

DI-Bench is a pipeline that automatically creates realistic data intelligence benchmarks for enterprise agents by linking data tables, dimensions, metrics, and documents into an artifact graph. It generates questions that combine structured data queries with knowledge retrieval, validates answers via query execution and LLM-generated questions, and has produced a 731-task benchmark covering knowledge retrieval, analytical computation, and rule‑grounded reasoning. Evaluation of four models on this benchmark shows that only 32% accuracy is achieved on computational tasks that involve business rules modifying the computation.

By Jiangyun Zhang, Kristen Surrao, Torpong Nitayanont, Yupei Zhang, Roopali Singh, Zhiyu Chen, Julia Huang, Zhou Tang, Shayan Ali Akbar, Omar Alonso, Erwin Cornejo, Yuan Li, Yi Zhang
arXiv AI
Sep 7

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

ERPBench is a benchmark that evaluates large language model agents in enterprise decision-making through a six‑round ERP simulation covering pricing, production, procurement, inventory, finance, and market competition. It tests the same 100 problems in two market ecologies—Solo, where agents compete against rule‑based opponents, and Arena, where six agents compete together—producing 1,200 model trajectories across 7,200 decision rounds. Results show that model performance varies by ecology, with DeepSeek best in Solo and Gemini best in Arena, and only 21 of 100 problems yield the same top performer across both settings.

By Xinran Zhang, Pengrui Lu, Lyumanshan Ye, Pengfei Liu
arXiv Machine Learning
Sep 1

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.

By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
arXiv AI
Aug 11

Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline

arXiv:2608. 09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema change, and queries that execute successfully while returning the wrong business number.

By Morris Lee