arXiv AI By So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

Read the original on arXiv AI →

arXiv:2607. 06482v1 Announce Type: cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 28

DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling

DataSTORM is an LLM‑based agentic system designed to conduct deep research over large‑scale structured databases and internet sources. It applies principles of Exploratory Data Analysis and Data Storytelling to frame research as a thesis‑driven analytical process, iteratively generating hypotheses, performing quantitative reasoning, and crafting coherent narratives. Evaluations on InsightBench and a new ACLED‑based dataset show that DataSTORM surpasses existing systems, achieving significant improvements in insight‑level recall and summary‑level scores.

By Shicheng Liu, Yucheng Jiang, Sajid Farook, Camila Nicollier Sanchez, David Fernando Castro Pena, Monica S. Lam
arXiv AI
Jun 10

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake

arXiv:2606. 10460v1 Announce Type: cross Abstract: Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved.

By Haonan Wang, Jiaxiang Liu, Yurong Liu, Austin Senna Wijaya, Tianle Zhou, Eden Wu, Yijia Chen, Wanting You, Reya Vir, Daniela Pinto, Grace Fan, Yusen Zhang, Juliana Freire, Eugene Wu
arXiv AI
Sep 10

Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?

The paper introduces Open Tabular Insight Extraction (OpenTI), a unified framework aimed at democratizing access to insights from large table corpora. It highlights how current research is fragmented across domains like table QA, text‑to‑SQL, and data analysis agents, and shows that existing systems and benchmarks fall short of covering the full end‑to‑end scope of OpenTI. The authors propose a consolidated terminology, conduct a systematic review, and outline a research agenda for developing comprehensive OpenTI systems, evaluation methods, and interaction paradigms.

By Daniel Gomm, Maarten de Rijke, Madelon Hulsebos
arXiv AI
Aug 5

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

arXiv:2608. 03451v1 Announce Type: new Abstract: Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia.

By Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo