arXiv Machine Learning

Same Quantity, Different Answer: Numerical Representation Invariance in Language Models

The paper introduces a benchmark of 3,600 exact‑rational word problems and 8,600 prompts that test whether language models give the same canonical answer when the same quantity is expressed in different numeric forms (decimal, fraction, percentage, number word, scientific notation, or unit‑converted). After normalizing answer syntax, canonical accuracy is high (0.969–0.996), but correctness across equivalent representations drops to 0.848–0.981, revealing that many errors stem from the evaluator’s number grammar rather than the models’ reasoning. The study also finds that representation consensus does not outperform paraphrase consensus on low‑error subsets and that certain models (e.g., Mistral Small 4) exhibit systematic unit‑conversion errors.

arXiv Computation and Language
Sep 22

Euston: Training Away Mathematical Sycophancy Without Losing the Mathematics

Euston is an 8‑B parameter mathematical claim‑verification model that resists producing false derivations when presented with corrupted theorems. It was trained on 3,026 matched true/corrupted statement pairs generated by GraphSynth, a probabilistic factor‑graph generator, and fine‑tuned from DeepSeek‑R1‑8B using GRPO. On a balanced held‑out split, Euston’s balanced accuracy rose from 29.50 % to 63.75 %, and its discrimination gap improved from –0.5 % to +27.5 %, while maintaining comparable general mathematical ability and reducing response length and truncation rates.

By Zehua Cheng, Wei Dai, Jiahao Sun
arXiv Machine Learning
Aug 19

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.

By Ayoub Kirouane, Christos Petrocheilos
arXiv AI
Sep 1

SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization

arXiv:2608.29270v1 Announce Type: cross Abstract: Autoformalization translates informal mathematical theorems into code for proof assistants such as Lean. A central challenge is that current evaluati...

By Hojae Han, Jongyoon Kim, Sanghyuk Park, Dongwook Cheon, Myungjae Jeon, Sunjong Choi, Soonho Kong, Wonseok Heo, Seung-won Hwang, Donghoon Hyeon