Hugging Face Trending Papers

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

Read the original on Hugging Face Trending Papers →

The Era by Eon benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from a company’s data. In the original benchmark, top models answered 22–25 of 27 questions, barely distinguishing performance. The updated benchmark adds eight templates that rely on hidden facts not explicitly stated in any question or document, forcing agents to infer information from indirect data. Twelve agents were evaluated, with the best achieving 18 of 24 correct answers, while the hardest questions—requiring selection among similar records—were answered correctly only 1 out of 84 attempts across all agents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 25

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

The "Era by Eon" benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from generated company data. While top models can answer most questions, the benchmark introduces eight new templates that rely on hidden facts not explicitly stated in any document, making the task harder. Evaluation of 12 agents shows that only the best agent correctly answers 18 of 24 attempts, with many questions remaining largely unsolved.

By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary
arXiv AI
Sep 11

The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

The Era by Eon Benchmark is a new dataset for evaluating large language model agents that interact with enterprise tools. It constructs a complete fictional company with product simulators, internal databases, and benchmark questions, all generated from a shared entity graph to ensure consistency. Exact answer keys are computed from the generated records, allowing precise grading and validation of realism and adversarial robustness across 23 simulated companies.

By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary
arXiv AI
Sep 7

Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

The paper introduces Agentic Context Cracking, a technique that adaptively and speculatively structures unstructured data during the reasoning process of large language model agents. By creating a sub-agent that extracts useful structure from documents as they are opened, the method reduces the need to repeatedly read large files, cutting token usage by 53% on the FanOutQA benchmark while maintaining accuracy. Over time, more queries are answered using the accumulated structured data, approaching the efficiency of a database lookup.

By Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos
arXiv AI
2d ago

AX is the New AEO

arXiv:2609.34951v2 Announce Type: replace Abstract: In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training...

By Ido Finder, Assaf Elovic, Gad Shalev, Liad Yosef