arXiv AI By Akrin Zheng, Alexander Wu, Alaia Liu

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

Read the original on arXiv AI →

arXiv:2608. 10679v1 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 31

Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge

The paper introduces ElephantBench, a closed‑book knowledge probe with 1,094 multi‑account factual questions generated via an auditable graph‑based pipeline that pulls documents from a low‑exposure web corpus and identifies naturally occurring disagreements. Across 32 large language models, even the best model only recovers both divergent accounts on 52.4% of questions, and most models recall one account while omitting the other, indicating persistent epistemic myopia. The study shows that scaling model size and inference‑time reasoning improves recall but does not eliminate incompleteness, and that exposure imbalance in the corpus biases models toward the dominant account.

By Zhuoshi Pan, Junru Lu, Yan Qian, H. Vicky Zhao, Di Yin, Xing Sun
arXiv AI
Aug 28

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

CorporateBench (CB) is a large‑scale, human‑validated Q&A benchmark designed to evaluate large language models on enterprise‑scale document collections. It contains over 230,000 documents derived from four synthetically generated firms, each modeled with a temporally evolving knowledge base that ensures logical consistency across hundreds of thousands of documents. The benchmark tests LLMs on information extraction and knowledge‑base querying, revealing that performance degrades as input size approaches realistic corporate scales.

By Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang, Igor Labutov
arXiv AI
Sep 25

Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora

The paper introduces ingest‑time fact compilation, an architecture that preprocesses and compiles corpus data into self‑contained facts with resolved revisions, deletions, and source trust. By storing this compiled state, query‑time models can retrieve answers directly, avoiding costly reconstruction from raw passages. Experiments show that this approach reduces read cost per question by 12.89× and token usage by 21.6× while maintaining accuracy.

By Kyle Wild, Yusuke Takahashi, Asako Uraki
arXiv AI
Sep 10

DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents

DI-Bench is a pipeline that automatically creates realistic data intelligence benchmarks for enterprise agents by linking data tables, dimensions, metrics, and documents into an artifact graph. It generates questions that combine structured data queries with knowledge retrieval, validates answers via query execution and LLM-generated questions, and has produced a 731-task benchmark covering knowledge retrieval, analytical computation, and rule‑grounded reasoning. Evaluation of four models on this benchmark shows that only 32% accuracy is achieved on computational tasks that involve business rules modifying the computation.

By Jiangyun Zhang, Kristen Surrao, Torpong Nitayanont, Yupei Zhang, Roopali Singh, Zhiyu Chen, Julia Huang, Zhou Tang, Shayan Ali Akbar, Omar Alonso, Erwin Cornejo, Yuan Li, Yi Zhang