The "Era by Eon" benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from generated company data. While top models can answer most questions, the benchmark introduces eight new templates that rely on hidden facts not explicitly stated in any document, making the task harder. Evaluation of 12 agents shows that only the best agent correctly answers 18 of 24 attempts, with many questions remaining largely unsolved.
By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary
The Era by Eon Benchmark is a new dataset for evaluating large language model agents that interact with enterprise tools. It constructs a complete fictional company with product simulators, internal databases, and benchmark questions, all generated from a shared entity graph to ensure consistency. Exact answer keys are computed from the generated records, allowing precise grading and validation of realism and adversarial robustness across 23 simulated companies.
By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary
Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded reasoning: locating sparse evidence across a large, messy collection of workplace files, reconciling inconsistent terminology, units, and time conventions, and computing an answer.
The paper introduces Agentic Context Cracking, a technique that adaptively and speculatively structures unstructured data during the reasoning process of large language model agents. By creating a sub-agent that extracts useful structure from documents as they are opened, the method reduces the need to repeatedly read large files, cutting token usage by 53% on the FanOutQA benchmark while maintaining accuracy. Over time, more queries are answered using the accumulated structured data, approaching the efficiency of a database lookup.
By Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos
arXiv:2609.34951v2 Announce Type: replace
Abstract: In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training...
By Ido Finder, Assaf Elovic, Gad Shalev, Liad Yosef
arXiv:2608.31082v1 Announce Type: new
Abstract: Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI...
By Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos
The paper introduces a method for continual enterprise world model discovery, enabling an agent to learn and adapt to business rules in dynamic systems without prior knowledge. Using a ServiceNow environment called EnterpriseWorldShift, the authors evaluate their Continual Discovery Agent (CDA) across four rule-modification scenarios—discovery, revision, extension, and retirement—showing that CDA predicts rule effects more accurately than lookup-based approaches, improving IoU by up to 8.98 points. The agent can answer queries from its internal model without querying the live system.
By Shambhavi Mishra, David Vazquez, Perouz Taslakian, Marco Pedersoli, Jose Dolz, Issam H. Laradji
DI-Bench is a pipeline that automatically creates realistic data intelligence benchmarks for enterprise agents by linking data tables, dimensions, metrics, and documents into an artifact graph. It generates questions that combine structured data queries with knowledge retrieval, validates answers via query execution and LLM-generated questions, and has produced a 731-task benchmark covering knowledge retrieval, analytical computation, and rule‑grounded reasoning. Evaluation of four models on this benchmark shows that only 32% accuracy is achieved on computational tasks that involve business rules modifying the computation.
By Jiangyun Zhang, Kristen Surrao, Torpong Nitayanont, Yupei Zhang, Roopali Singh, Zhiyu Chen, Julia Huang, Zhou Tang, Shayan Ali Akbar, Omar Alonso, Erwin Cornejo, Yuan Li, Yi Zhang
arXiv:2607. 12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy.
By Ning Liu
Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across heterogeneous sources.
arXiv:2605.23916v2 Announce Type: replace-cross
Abstract: AI agents often pick tools from registries, where each tool's provider writes its description. We ask whether sales language in those descrip...
By Haochuan Kevin Wang, Zechen Zhang
arXiv:2608. 10679v1 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers.
By Akrin Zheng, Alexander Wu, Alaia Liu