arXiv AI By Chen Wang, Junzhe Zhao, Xin Cong, Wanlu Deng, Ke Deng

Benchmarking Language Models for Statistical Problem Formulation

Read the original on arXiv AI →

Large language models are increasingly used to assist with statistical and data science tasks, but current evaluations assume the analysis goal is already defined. This paper formalizes the upstream step of Statistical Problem Formulation into two subtasks—classification of the statistical problem and identification of relevant variables—and introduces StatFormBench, a benchmark comprising 1,013 samples from five statistics textbooks and a data science case library. Across 14 open- and closed‑source LLMs, the best zero‑shot models achieve only 72.0% fine‑grained classification accuracy and 63.2% variable set overlap, with no model consistently excelling in both subtasks and limited gains from enhanced prompting strategies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 9

TinyJudge: Unverifiable Constraint Alignment via Lightweight Specialist Ensembles

arXiv:2606. 07520v1 Announce Type: cross Abstract: Instruction Following (IF) is a core capability of LLMs, requiring strict adherence to diverse constraints, ranging from verifiable ones (e.

By Yirong Zeng, Yufei Liu, Xiao Ding, Yutai Hou, Yuxian Wang, Wu Ning, Haonan Song, Dandan Tu, Qixun Zhang, Yuxiang He, Bibo Cai, Ting Liu
arXiv AI
Jul 8

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

arXiv:2607. 05992v1 Announce Type: cross Abstract: Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites.

By Daryna Dementieva, Nikolay Babakov, Kathy H\"ammerl, Ilseyar Alimova, Jind\v{r}ich Libovick\'y, Shu Okabe, Miras Baisbay, Lukas Edman, Abrorkhon Inomkhujaev, Antonia Karamolegkou, Mateusz Lango, Volkan \"Ozer, Nikola Selic, Subhankar Swain, Tsedeniya Kinfe Temesgen, Galit Bary Weisberg, Alexander Fraser