arXiv AI

Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity

arXiv:2608. 16118v1 Announce Type: new Abstract: How should we assess whether large language models can perform mathematical invention?

arXiv Machine Learning
Aug 27

Emergent Abilities in Large Language Models: A Survey

Emergent Abilities in Large Language Models: A Survey reviews how scaling LLMs leads to previously unseen capabilities such as advanced reasoning, in-context learning, coding, and problem-solving. The paper critically examines definitions, inconsistencies, and the conditions that foster these abilities, including scaling laws, task complexity, pre‑training loss, quantization, and prompting strategies. It also discusses the extension to Large Reasoning Models and highlights safety concerns like deception, manipulation, and reward hacking, calling for improved evaluation and governance.

By Leonardo Berti, Flavio Giorgi, Gjergji Kasneci
arXiv AI
Aug 14

Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design

arXiv:2603. 08322v2 Announce Type: replace Abstract: We study mathematical discovery through the lens of neurosymbolic reasoning, where an AI agent powered by a large language model (LLM), coupled with symbolic computation tools, and human strategic direction, jointly produced a new result in combinatorial design theory.

By Hai Xia, Carla P. Gomes, Bart Selman, Stefan Szeider
arXiv AI
Sep 3

AI Mathematician: Towards Fully Automated Frontier Mathematical Research

The paper introduces the AI Mathematician (AIM) framework, which leverages Large Reasoning Models (LRMs) to tackle frontier mathematical research. AIM addresses the complexity and procedural rigor of research problems through an exploration mechanism for longer solution paths and a pessimistic reasonable verification method for reliability. Early experiments show AIM can autonomously construct significant proof components and uncover non‑trivial insights across real‑world mathematical topics.

By Yuanhang Liu, Yanxing Huang, Yanqiao Wang, Peng Li, Yang Liu
arXiv Machine Learning
Sep 24

Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models

The paper investigates how large language models (LLMs) balance capability and efficiency when using Chain-of-Thought reasoning on arithmetic and algorithmic tasks. It finds that while larger models solve more problems correctly, the improvement follows an exponential decay that slows with scale, indicating diminishing returns. Additionally, the length of reasoning output grows with problem size but does not improve with larger models, suggesting efficiency does not benefit from scaling.

By Moritz Laber, Zohair Shafi, Germans Savcisens, Brennan Klein, Matteo Chinazzi, Samuel V. Scarpino, Albert-L\'aszl\'o Barab\'asi, Tina Eliassi-Rad
Hugging Face Trending Papers
Jul 6

Knowledge Knows, Verbalization Tells: Disentangling Latent Directions for Mathematical Solvability in LLMs

Although LLMs have made significant progress in mathematical reasoning, determining whether a mathematical problem is solvable remains a fundamental yet challenging capability. While recent studies have probed internal representations of model solvability beliefs, verbalization has primarily been studied behaviorally rather than as an internal representation, limiting its analysis and manipulation.