arXiv:2506. 12339v2 Announce Type: replace-cross Abstract: We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural language instructions.
By Xi Cheng, Ruiyan Zhu, Ke Liu, Rakesh Chowdary Machineni, Lyuhao Chen, Brian Zhu, Daniel Jin, Zheng Qi, Neeraj Parihar, Zhoutian Xu, Oliver Gao
Spreadsheets and tables are widely used representations for structured data analysis, but effective analysis still requires substantial manual effort and domain expertise. Recent large language model (LLM) agents can automate parts of this process, but they often provide limited transparency into intermediate decisions, rely on implicit assumptions, struggle with multi-table comparison, and repeat similar workflows without adapting to a user's preferences.
The paper introduces Table Graph Reasoner (TabGR), a model that represents tables as an Attributed Table Graph (ATG) to preserve row-column-cell structure and enable graph-based reasoning without task-specific training. It also proposes a Question-Guided Personalized PageRank (QG-PPR) mechanism to rerank tabular data and address the lost-in-the-middle issue. Experiments on multiple table reasoning benchmarks show that TabGR outperforms state-of-the-art models by up to 9.7% in accuracy.
By Yuxiang Wang, Junhao Gan, Shengxiang Gao, Shenghao Ye, Zhengyi Yang, Jianzhong Qi
Tables are ubiquitous across diverse domains, yet reasoning over them remains a significant challenge for modern large language models (LLMs). Current approaches typically linearize tables into sequen...
The paper proposes a new framework that improves spreadsheet chunking for large language model (LLM)-driven retrieval-augmented generation (RAG) systems by adding semantic cell annotations. This approach outperforms current state‑of‑the‑art methods but is limited by the inherent two‑dimensional, unstructured nature of spreadsheets, which cannot be fully captured by finite classification categories. The authors argue that future progress requires dimensionality‑reduction techniques to flatten spreadsheets into one‑dimensional text, simplifying downstream RAG interpretation and generation.
By Zofia Smole\'n
The paper proposes a new framework that uses semantic cell annotation to split spreadsheets into interpretable chunks for large language model (LLM)-driven Retrieval-Augmented Generation (RAG) systems. This approach improves answer generation by providing richer context rather than merely enhancing retrieval accuracy. However, the authors argue that the inherent two‑dimensional, unstructured nature of spreadsheets imposes a hard ceiling on classification‑based methods, suggesting that future work should focus on dimensionality‑reduction techniques to flatten spreadsheets into one‑dimensional text for easier processing by RAG.