Spreadsheets are a primary medium for publishing tabular data, yet automatically extracting structured content from them remains difficult due to heterogeneous layouts, diverse file formats, and inconsistent organizational conventions. We address two core tasks in spreadsheet understanding: Cell-Type Classification (CTC), which assigns roles to cells, and Table Detection (TD), which identifies table bounding boxes within sheets.
arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.
By Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzm\"uller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Ga\"el Varoquaux, Frank Hutter
The paper proposes a new framework that improves spreadsheet chunking for large language model (LLM)-driven retrieval-augmented generation (RAG) systems by adding semantic cell annotations. This approach outperforms current state‑of‑the‑art methods but is limited by the inherent two‑dimensional, unstructured nature of spreadsheets, which cannot be fully captured by finite classification categories. The authors argue that future progress requires dimensionality‑reduction techniques to flatten spreadsheets into one‑dimensional text, simplifying downstream RAG interpretation and generation.
By Zofia Smole\'n
arXiv:2605. 05811v2 Announce Type: replace Abstract: Workbook-scale spreadsheet understanding is increasingly important for language-model-based data analysis agents, but remains challenging because relevant information is often distributed across multiple sheets with heterogeneous schemas, layouts, and implicit relationships.
By Yiming Lei, Yuhang Yao, Yujia Zhang, Yiqi Wang, Bo Guan, Depei Zhu, Chunhui Wang, Zhuonan Hao, Tianyu Shi
The paper proposes a new framework that uses semantic cell annotation to split spreadsheets into interpretable chunks for large language model (LLM)-driven Retrieval-Augmented Generation (RAG) systems. This approach improves answer generation by providing richer context rather than merely enhancing retrieval accuracy. However, the authors argue that the inherent two‑dimensional, unstructured nature of spreadsheets imposes a hard ceiling on classification‑based methods, suggesting that future work should focus on dimensionality‑reduction techniques to flatten spreadsheets into one‑dimensional text for easier processing by RAG.
arXiv:2606. 29955v1 Announce Type: cross Abstract: Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making.
By Jian Zhu, Yuzheng Zhang, Zeyao Ma, Bohan Zhang, Armin Schoepf, Daniel Woloch, Peter Yiliu Wang, Guangyu Robert Yang, Samuel Jacob, Siddharth Nagisetty, Abhiram Chundru, Jean Lin, Spencer Mateega, Jing Zhang