The paper investigates whether large language models (LLMs) possess coherent, human-like knowledge structures in mathematical reasoning by applying Knowledge Space Theory (KST). Using a KST-based framework, the authors evaluate eight open- and closed-source LLMs and find that they frequently violate knowledge dependencies, fail to leverage related context, and exhibit low overlap in knowledge distributions compared to real human learners. These structural deficiencies remain largely invisible to standard accuracy or LLM-as-judge evaluations, suggesting that current LLMs do not follow a human-like knowledge structure.
By Peng Cui, Heejin Do, Mrinmaya Sachan
And an Overview of Recent Inference-Scaling Papers
By Sebastian Raschka, PhD
The paper introduces T2T (Thickening-to-Thinning), a dynamic reward framework for large language models that mimics human learning by separating exploration and consolidation phases. During incorrect attempts, T2T encourages exploration to broaden the search space, while after correct solutions it applies length penalties to promote concise reasoning. Experiments on mathematical benchmarks across five mainstream LLMs show that T2T outperforms standard GRPO and recent baselines, improving overall reasoning performance.
By Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang
arXiv:2410.23912v3 Announce Type: replace-cross
Abstract: The reasoning abilities of large language models (LLMs) have improved with chain-of-thought (CoT) prompting, allowing models to solve complex...
By Fu-Chieh Chang, Yu-Ting Lee, Hui-Ying Shih, Yi Hsuan Tseng, Pei-Yuan Wu
arXiv:2607. 21856v1 Announce Type: new Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs.
By Ziran Yang, Chengshuai Shi, Raj Ghugare, Benjamin Eysenbach, Karthik Narasimhan, Chi Jin