arXiv AI

SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation

arXiv:2506. 12339v2 Announce Type: replace-cross Abstract: We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural language instructions.

arXiv AI
Jun 30

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

arXiv:2606. 29955v1 Announce Type: cross Abstract: Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making.

By Jian Zhu, Yuzheng Zhang, Zeyao Ma, Bohan Zhang, Armin Schoepf, Daniel Woloch, Peter Yiliu Wang, Guangyu Robert Yang, Samuel Jacob, Siddharth Nagisetty, Abhiram Chundru, Jean Lin, Spencer Mateega, Jing Zhang
Hugging Face Trending Papers
Jun 9

TabClaw: An Interactive and Self-Evolving Agent for Spreadsheet Manipulation and Table Reasoning

Spreadsheets and tables are widely used representations for structured data analysis, but effective analysis still requires substantial manual effort and domain expertise. Recent large language model (LLM) agents can automate parts of this process, but they often provide limited transparency into intermediate decisions, rely on implicit assumptions, struggle with multi-table comparison, and repeat similar workflows without adapting to a user's preferences.

arXiv AI
4d ago

AnyAct: Universal Action for Self-Evolving Agents

AnyAct introduces a universal action layer that consolidates diverse tool capabilities into a self‑evolving action space for AI agents operating in open‑world environments. It tackles the scale dilemma, tool non‑stationarity, and heterogeneous feedback by using hierarchical progressive retrieval and test‑time reliability evolution, while a heterogeneous observation grounding module unifies multi‑modal feedback. Evaluations on LiveMCPBench and the newly created OSMCP benchmark show state‑of‑the‑art performance, with significant gains in task success rate and reduced execution steps, especially for models with limited native capabilities.

By Lingrui Xu, Yangqin Jiang, Jiachang Zhang, Xubin Ren, Chao Huang
arXiv AI
Aug 11

Back to the Future: A workbook time machine for spread sheet creation benchmarks

arXiv:2608. 07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting).

By Mansi Uniyal, Agamdeep Singh, Ananya Singha, Priyanshu Gupta, Mukul Singh, Gust Verbruggen, Vu Le, Sumit Gulwani
arXiv AI
Jul 7

SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks

arXiv:2603. 10002v2 Announce Type: replace-cross Abstract: We consider the task of end-to-end spreadsheet generation, where language models produce spreadsheet artifacts to satisfy users' explicit and implicit constraints, specified in natural language.

By Srivatsa Kundurthy, Clara Na, Michael Handley, Zach Kirshner, Chen Bo Calvin Zhang, Manasi Sharma, Emma Strubell, John Ling
arXiv AI
6d ago

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.

By Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasovi\'c