The paper introduces a synthetic data generator that creates fully consistent, fictional enterprises—complete with workforce, customers, sales, support, and communication records—without relying on any real dataset. It validates realism through a five‑axis scorecard, an adversarial detector, and soundness checks, achieving a mean realism score of 99.1 across 23 generated companies. A second generator produces relational databases from business questions, ensuring qualifying rows and exact labels, and is available as a hosted service and containerized simulators.
By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary, Omer Niv
The "Era by Eon" benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from generated company data. While top models can answer most questions, the benchmark introduces eight new templates that rely on hidden facts not explicitly stated in any document, making the task harder. Evaluation of 12 agents shows that only the best agent correctly answers 18 of 24 attempts, with many questions remaining largely unsolved.
By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary
arXiv:2609.37658v1 Announce Type: new
Abstract: LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-te...
By Min Yang, Yichen Pan, Jinghua Piao, Dandan Song, Yongshun Gong, Yong Li
DI-Bench is a pipeline that automatically creates realistic data intelligence benchmarks for enterprise agents by linking data tables, dimensions, metrics, and documents into an artifact graph. It generates questions that combine structured data queries with knowledge retrieval, validates answers via query execution and LLM-generated questions, and has produced a 731-task benchmark covering knowledge retrieval, analytical computation, and rule‑grounded reasoning. Evaluation of four models on this benchmark shows that only 32% accuracy is achieved on computational tasks that involve business rules modifying the computation.
By Jiangyun Zhang, Kristen Surrao, Torpong Nitayanont, Yupei Zhang, Roopali Singh, Zhiyu Chen, Julia Huang, Zhou Tang, Shayan Ali Akbar, Omar Alonso, Erwin Cornejo, Yuan Li, Yi Zhang
arXiv:2608. 10679v1 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers.
By Akrin Zheng, Alexander Wu, Alaia Liu
The Era by Eon benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from a company’s data. In the original benchmark, top models answered 22–25 of 27 questions, barely distinguishing performance. The updated benchmark adds eight templates that rely on hidden facts not explicitly stated in any question or document, forcing agents to infer information from indirect data. Twelve agents were evaluated, with the best achieving 18 of 24 correct answers, while the hardest questions—requiring selection among similar records—were answered correctly only 1 out of 84 attempts across all agents.