arXiv:2608. 09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema change, and queries that execute successfully while returning the wrong business number.
By Morris Lee
The Era by Eon Benchmark is a new dataset for evaluating large language model agents that interact with enterprise tools. It constructs a complete fictional company with product simulators, internal databases, and benchmark questions, all generated from a shared entity graph to ensure consistency. Exact answer keys are computed from the generated records, allowing precise grading and validation of realism and adversarial robustness across 23 simulated companies.
By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary
arXiv:2608. 07946v1 Announce Type: cross Abstract: Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean.
By Mike Helwig
The "Era by Eon" benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from generated company data. While top models can answer most questions, the benchmark introduces eight new templates that rely on hidden facts not explicitly stated in any document, making the task harder. Evaluation of 12 agents shows that only the best agent correctly answers 18 of 24 attempts, with many questions remaining largely unsolved.
By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary
The Era by Eon benchmark tests enterprise agents by presenting questions that specify answer rules and require code to compute answers from a company’s data. In the original benchmark, top models answered 22–25 of 27 questions, barely distinguishing performance. The updated benchmark adds eight templates that rely on hidden facts not explicitly stated in any question or document, forcing agents to infer information from indirect data. Twelve agents were evaluated, with the best achieving 18 of 24 correct answers, while the hardest questions—requiring selection among similar records—were answered correctly only 1 out of 84 attempts across all agents.
A practical walkthrough using text-to-SQL as the example The post Why I Stopped Using One Agent and Built a Multi-Agent Pipeline Instead appeared first on Towards Data Science .
By Priyansh Bhardwaj
arXiv:2609.22917v1 Announce Type: new
Abstract: Non-technical stakeholders frequently cannot write the SQL needed to extract insights from operational databases. We built and evaluated a Text-to-SQL...
By Vigneshwar Ravi Rao, Rupesh Swarnakar, Fayeq Jeelani Syed{\dag}
The paper reports a case study of a large language model (LLM) coding agent tasked with building a multi‑component data system from a detailed specification. During a single session the agent introduced five defects, which were categorized by violated constraints and detection methods. The study also evaluates the agent’s retrieval‑filtering strategy on the HotpotQA benchmark, showing that filtering to a graph‑identified entity set yields higher recall than unfiltered search, with a statistically significant gap across all tested budgets.
By Phanindra Reddy Madduru
Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integri...
ContractEval is a diagnostic framework that makes active obligations in procedural instructions explicit by representing them as query‑conditioned obligations. It matches these obligations against response or trace evidence, identifying omissions, wrong branches, ordering errors, extra actions, invariant breaches, and output‑contract violations as distinct conformance failures. In tests on audited procedural contracts, ContractEval detects and localizes all injected structural failures that output‑only and trace‑aware LLM judges miss, though it is not a compliance guarantee and remains calibration‑sensitive.
By Praphul Singh, Shanu Kumar, Akshat Agarwal, Ganesh Kumar
DNative‑Twin is a graph‑native digital twin that records an AI agent’s committed decision as a typed trajectory, linking observed state, decision path, and authority. It re‑executes the decision mechanism under declared conditions, synchronizing and replaying the process in isolation to compare outcomes under controlled changes. Experiments on enterprise decision logs show that adding replay‑contract state and verification evidence improves unresolved‑divergence recall from 0 to 1.0, while end‑to‑end processing time rises from 0.794 to 8.889 seconds across 500–5,000 cases.
By Junjie Pang, Zhenzhen Xie, Haoke Han, Ying He, Jing Wang, Gang Liu
arXiv:2607. 07397v1 Announce Type: new Abstract: Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs.
By Elaine Ang, Chenxi Huang, Georgios Liargkovas, Jerry Liu, Jinhui Liu, Nikos Pagonas, Charlie Summers, Haonan Wang, Jiakai Xu, Tianle Zhou, Yusen Zhang, Zhou Yu, Zhuo Zhang, Tianyi Peng, Kostis Kaffes, Eugene Wu