arXiv AI By Dominika Agnieszka D{\l}ugosz, Arlindo Oliveira, Natalia D\'iaz-Rodr\'iguez

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

Read the original on arXiv AI →

The paper re‑examines the GSM‑Symbolic benchmark, arguing that its claim of widespread reasoning failures in 25 LLMs is based on weak statistics. Using bootstrapped Generalised Linear Mixed Models on 20 open‑weight models, only eight show significant performance changes, and a systematic shift toward larger integers in the dataset explains many of these effects. The authors also uncover model‑specific failure modes such as variable binding fragility, arithmetic limits, and dual‑task interference, cautioning against blanket conclusions about LLM reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 14

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Bangladesh, where over 230 million people speak Bengali.