arXiv AI By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary

The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

Read the original on arXiv AI →

The Era by Eon Benchmark is a new dataset for evaluating large language model agents that interact with enterprise tools. It constructs a complete fictional company with product simulators, internal databases, and benchmark questions, all generated from a shared entity graph to ensure consistency. Exact answer keys are computed from the generated records, allowing precise grading and validation of realism and adversarial robustness across 23 simulated companies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 12

Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data

The paper introduces a synthetic data generator that creates fully consistent, fictional enterprises—complete with workforce, customers, sales, support, and communication records—without relying on any real dataset. It validates realism through a five‑axis scorecard, an adversarial detector, and soundness checks, achieving a mean realism score of 99.1 across 23 generated companies. A second generator produces relational databases from business questions, ensuring qualifying rows and exact labels, and is available as a hosted service and containerized simulators.

By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary, Omer Niv
arXiv AI
Sep 25

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

The "Era by Eon" benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from generated company data. While top models can answer most questions, the benchmark introduces eight new templates that rely on hidden facts not explicitly stated in any document, making the task harder. Evaluation of 12 agents shows that only the best agent correctly answers 18 of 24 attempts, with many questions remaining largely unsolved.

By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary
arXiv AI
Sep 10

DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents

DI-Bench is a pipeline that automatically creates realistic data intelligence benchmarks for enterprise agents by linking data tables, dimensions, metrics, and documents into an artifact graph. It generates questions that combine structured data queries with knowledge retrieval, validates answers via query execution and LLM-generated questions, and has produced a 731-task benchmark covering knowledge retrieval, analytical computation, and rule‑grounded reasoning. Evaluation of four models on this benchmark shows that only 32% accuracy is achieved on computational tasks that involve business rules modifying the computation.

By Jiangyun Zhang, Kristen Surrao, Torpong Nitayanont, Yupei Zhang, Roopali Singh, Zhiyu Chen, Julia Huang, Zhou Tang, Shayan Ali Akbar, Omar Alonso, Erwin Cornejo, Yuan Li, Yi Zhang
Hugging Face Trending Papers
Sep 24

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

The Era by Eon benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from a company’s data. In the original benchmark, top models answered 22–25 of 27 questions, barely distinguishing performance. The updated benchmark adds eight templates that rely on hidden facts not explicitly stated in any question or document, forcing agents to infer information from indirect data. Twelve agents were evaluated, with the best achieving 18 of 24 correct answers, while the hardest questions—requiring selection among similar records—were answered correctly only 1 out of 84 attempts across all agents.