Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across heterogeneous sources.
The paper introduces ElephantBench, a closed‑book knowledge probe with 1,094 multi‑account factual questions generated via an auditable graph‑based pipeline that pulls documents from a low‑exposure web corpus and identifies naturally occurring disagreements. Across 32 large language models, even the best model only recovers both divergent accounts on 52.4% of questions, and most models recall one account while omitting the other, indicating persistent epistemic myopia. The study shows that scaling model size and inference‑time reasoning improves recall but does not eliminate incompleteness, and that exposure imbalance in the corpus biases models toward the dominant account.
By Zhuoshi Pan, Junru Lu, Yan Qian, H. Vicky Zhao, Di Yin, Xing Sun
CorporateBench (CB) is a large‑scale, human‑validated Q&A benchmark designed to evaluate large language models on enterprise‑scale document collections. It contains over 230,000 documents derived from four synthetically generated firms, each modeled with a temporally evolving knowledge base that ensures logical consistency across hundreds of thousands of documents. The benchmark tests LLMs on information extraction and knowledge‑base querying, revealing that performance degrades as input size approaches realistic corporate scales.
By Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang, Igor Labutov
arXiv:2609.08869v1 Announce Type: cross
Abstract: Analysts in emerging equity markets keep answering the same questions. Did fundamentals match the market's response? How does the local currency co-m...
By Furqan Nasir, Muhammad Atif Saeed, Muhammad Ehsan, Sher Jeel Ahmad, Abdul Moiz Altaf
The paper introduces ingest‑time fact compilation, an architecture that preprocesses and compiles corpus data into self‑contained facts with resolved revisions, deletions, and source trust. By storing this compiled state, query‑time models can retrieve answers directly, avoiding costly reconstruction from raw passages. Experiments show that this approach reduces read cost per question by 12.89× and token usage by 21.6× while maintaining accuracy.
By Kyle Wild, Yusuke Takahashi, Asako Uraki
DI-Bench is a pipeline that automatically creates realistic data intelligence benchmarks for enterprise agents by linking data tables, dimensions, metrics, and documents into an artifact graph. It generates questions that combine structured data queries with knowledge retrieval, validates answers via query execution and LLM-generated questions, and has produced a 731-task benchmark covering knowledge retrieval, analytical computation, and rule‑grounded reasoning. Evaluation of four models on this benchmark shows that only 32% accuracy is achieved on computational tasks that involve business rules modifying the computation.
By Jiangyun Zhang, Kristen Surrao, Torpong Nitayanont, Yupei Zhang, Roopali Singh, Zhiyu Chen, Julia Huang, Zhou Tang, Shayan Ali Akbar, Omar Alonso, Erwin Cornejo, Yuan Li, Yi Zhang