Hugging Face Trending Papers

Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure

Read the original on Hugging Face Trending Papers →

The paper proposes a new framework that uses semantic cell annotation to split spreadsheets into interpretable chunks for large language model (LLM)-driven Retrieval-Augmented Generation (RAG) systems. This approach improves answer generation by providing richer context rather than merely enhancing retrieval accuracy. However, the authors argue that the inherent two‑dimensional, unstructured nature of spreadsheets imposes a hard ceiling on classification‑based methods, suggesting that future work should focus on dimensionality‑reduction techniques to flatten spreadsheets into one‑dimensional text for easier processing by RAG.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 18

Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure

The paper proposes a new framework that improves spreadsheet chunking for large language model (LLM)-driven retrieval-augmented generation (RAG) systems by adding semantic cell annotations. This approach outperforms current state‑of‑the‑art methods but is limited by the inherent two‑dimensional, unstructured nature of spreadsheets, which cannot be fully captured by finite classification categories. The authors argue that future progress requires dimensionality‑reduction techniques to flatten spreadsheets into one‑dimensional text, simplifying downstream RAG interpretation and generation.

By Zofia Smole\'n
Hugging Face Trending Papers
Aug 17

Structured Prediction for Scalable Spreadsheet Table Understanding: From Cell Types to Table Ranges (Extended Version)

Spreadsheets are a primary medium for publishing tabular data, yet automatically extracting structured content from them remains difficult due to heterogeneous layouts, diverse file formats, and inconsistent organizational conventions. We address two core tasks in spreadsheet understanding: Cell-Type Classification (CTC), which assigns roles to cells, and Table Detection (TD), which identifies table bounding boxes within sheets.

arXiv Machine Learning
Aug 18

Structured Prediction for Scalable Spreadsheet Table Understanding: From Cell Types to Table Ranges (Extended Version)

arXiv:2608. 16050v1 Announce Type: cross Abstract: Spreadsheets are a primary medium for publishing tabular data, yet automatically extracting structured content from them remains difficult due to heterogeneous layouts, diverse file formats, and inconsistent organizational conventions.

By Antoine Gauquier, Ioana Manolescu, Pierre Senellart
arXiv AI
Jul 7

SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks

arXiv:2603. 10002v2 Announce Type: replace-cross Abstract: We consider the task of end-to-end spreadsheet generation, where language models produce spreadsheet artifacts to satisfy users' explicit and implicit constraints, specified in natural language.

By Srivatsa Kundurthy, Clara Na, Michael Handley, Zach Kirshner, Chen Bo Calvin Zhang, Manasi Sharma, Emma Strubell, John Ling
arXiv AI
Jul 29

Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding

arXiv:2605. 05811v2 Announce Type: replace Abstract: Workbook-scale spreadsheet understanding is increasingly important for language-model-based data analysis agents, but remains challenging because relevant information is often distributed across multiple sheets with heterogeneous schemas, layouts, and implicit relationships.

By Yiming Lei, Yuhang Yao, Yujia Zhang, Yiqi Wang, Bo Guan, Depei Zhu, Chunhui Wang, Zhuonan Hao, Tianyu Shi