Same Quantity, Different Answer: Numerical Representation Invariance in Language Models
Read the original on arXiv Machine Learning →The paper introduces a benchmark of 3,600 exact‑rational word problems and 8,600 prompts that test whether language models give the same canonical answer when the same quantity is expressed in different numeric forms (decimal, fraction, percentage, number word, scientific notation, or unit‑converted). After normalizing answer syntax, canonical accuracy is high (0.969–0.996), but correctness across equivalent representations drops to 0.848–0.981, revealing that many errors stem from the evaluator’s number grammar rather than the models’ reasoning. The study also finds that representation consensus does not outperform paraphrase consensus on low‑error subsets and that certain models (e.g., Mistral Small 4) exhibit systematic unit‑conversion errors.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.