arXiv AI By Yongfu Zhu, Lin Sun, Jinzhu Wu, Weihong Lin, Xiaoqi Jian, Guangxiang Zhao, Change Jia, Linglin Zhang, Sai-er Hu, Yuhan Wu, Xiangzheng Zhang

Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design

Read the original on arXiv AI →

The paper examines how evaluation design can cause significant fluctuations in the benchmark results of reasoning models, particularly the Deepseek‑R1‑Distill series. It shows that subtle changes in evaluation conditions lead to large variations in reported performance, a phenomenon also seen in other open‑source models fine‑tuned from Deepseek‑R1‑Distill and in the QwQ‑32B model. The authors call for a more rigorous evaluation paradigm and provide empirical assessments of the Deepseek‑R1‑Distill models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 24

Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models

The paper investigates how large language models (LLMs) balance capability and efficiency when using Chain-of-Thought reasoning on arithmetic and algorithmic tasks. It finds that while larger models solve more problems correctly, the improvement follows an exponential decay that slows with scale, indicating diminishing returns. Additionally, the length of reasoning output grows with problem size but does not improve with larger models, suggesting efficiency does not benefit from scaling.

By Moritz Laber, Zohair Shafi, Germans Savcisens, Brennan Klein, Matteo Chinazzi, Samuel V. Scarpino, Albert-L\'aszl\'o Barab\'asi, Tina Eliassi-Rad