OpenAI Blog

Solving math word problems

We’ve trained a system that solves grade school math problems with nearly twice the accuracy of a fine-tuned GPT-3 model. It solves about 90% as many problems as real kids: a small sample of 9-12 year olds scored 60% on a test from our dataset, while our system scored 55% on those same problems.

arXiv AI
Sep 10

Tracing Mathematical Proficiency Through Problem-Solving Processes

The paper introduces Knowledge Tracing Leveraging Problem‑Solving Process (KT‑PSP), a method that incorporates students’ problem‑solving steps to model mathematical proficiency more comprehensively than traditional knowledge tracing. It presents the KT‑PSP‑25 dataset and a new framework, StatusKT, which uses a teacher‑student‑teacher LLM pipeline to extract proficiency indicators, generate responses, and evaluate mastery. Experiments show that StatusKT improves prediction accuracy and offers interpretable explanations by explicitly modeling proficiency.

By Jungyang Park, Suho Kang, Jaewoo Park, Jaehong Kim, Jaewoo Shin, Seonjoon Park, Youngjae Yu
arXiv AI
Sep 17

A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning

The paper presents a mechanistic analysis of how large language models solve grade‑school math word problems. It identifies a four‑stage sequential pipeline—Schema Abstraction, Operation Planning, Operand Binding, and Computation—each represented in distinct layer bands. The study shows that inserting an irrelevant clause disrupts the Operation Planning stage, pinpointing the cause of failure to specific attention heads.

By Zhongdi Qu, Carla P. Gomes
OpenAI Blog
Feb 2, 2022

Solving (some) formal math olympiad problems

We built a neural theorem prover for Lean that learned to solve a variety of challenging high-school olympiad problems, including problems from the AMC12 and AIME competitions, as well as two problems adapted from the IMO.

arXiv AI
Jun 3

PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models

arXiv:2606. 03858v1 Announce Type: new Abstract: Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by integrating numerical processing and mathematical reasoning, hindering the interpretability of failures in math tasks.

By Zetian Ouyang, Linlin Wang, Gerard de Melo, Liang He
arXiv AI
Jul 21

Probing the Difficulty Perception Mechanism of Large Language Models

arXiv:2510. 05969v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed on complex reasoning tasks, yet little is known about their ability to internally evaluate problem difficulty, which is an essential capability for adaptive reasoning and efficient resource allocation.

By Sunbowen Lee, Qingyu Yin, Chak Tou Leong, Jialiang Zhang, Yicheng Gong, Shiwen Ni, Min Yang, Xiaoyu Shen