The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic
Read the original on arXiv AI →The paper re‑examines the GSM‑Symbolic benchmark, arguing that its claim of widespread reasoning failures in 25 LLMs is based on weak statistics. Using bootstrapped Generalised Linear Mixed Models on 20 open‑weight models, only eight show significant performance changes, and a systematic shift toward larger integers in the dataset explains many of these effects. The authors also uncover model‑specific failure modes such as variable binding fragility, arithmetic limits, and dual‑task interference, cautioning against blanket conclusions about LLM reasoning.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.