Benchmarking Language Models for Statistical Problem Formulation
Read the original on arXiv AI →Large language models are increasingly used to assist with statistical and data science tasks, but current evaluations assume the analysis goal is already defined. This paper formalizes the upstream step of Statistical Problem Formulation into two subtasks—classification of the statistical problem and identification of relevant variables—and introduces StatFormBench, a benchmark comprising 1,013 samples from five statistics textbooks and a data science case library. Across 14 open- and closed‑source LLMs, the best zero‑shot models achieve only 72.0% fine‑grained classification accuracy and 63.2% variable set overlap, with no model consistently excelling in both subtasks and limited gains from enhanced prompting strategies.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.