arXiv AI

Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory

The paper investigates whether large language models (LLMs) possess coherent, human-like knowledge structures in mathematical reasoning by applying Knowledge Space Theory (KST). Using a KST-based framework, the authors evaluate eight open- and closed-source LLMs and find that they frequently violate knowledge dependencies, fail to leverage related context, and exhibit low overlap in knowledge distributions compared to real human learners. These structural deficiencies remain largely invisible to standard accuracy or LLM-as-judge evaluations, suggesting that current LLMs do not follow a human-like knowledge structure.

arXiv AI
Jun 2

KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models

arXiv:2604. 17621v2 Announce Type: replace Abstract: Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based reasoning over that universe, a phenomenon we term "the tip of the iceberg.

By Xiao Zhang, Qianru Meng, Yongjian Chen, Yumeng Wang, Johan Bos
arXiv Machine Learning
3d ago

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

arXiv:2610. 02191v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions.

By Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez, Mohit Bansal, Zhengzhong Tu
arXiv AI
Aug 19

Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning

The paper introduces the Structure-Internalized Rule Language Model (SIRLM) to improve Knowledge Graph Reasoning (KGR) by addressing the mismatch between KG structural context and Large Language Model (LLM) parametric knowledge. SIRLM centers on a Structure-Internalized Rule Generator (SIRG) that uses in-context learning, a structural relation memory, a KG tokenizer, and a neuro-symbolic reasoner to generate structural rules and provide faithful rule-execution feedback. Experiments on 36 datasets against 17 state‑of‑the‑art KGR methods show that SIRLM achieves significant performance gains.

By Xingrui Zhuo, Jiapu Wang, Manzong Huang, Gongqing Wu, Xindong Wu
arXiv AI
Sep 25

LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches

LiveMathematicianBench is a dynamic multiple‑choice benchmark for research‑level mathematical reasoning, built from recent arXiv papers published after model training cutoffs. It introduces a thirteen‑category logical taxonomy of theorem types and uses a proof‑sketch‑guided distractor pipeline to create plausible but invalid answer choices, enhancing sensitivity to genuine reasoning. Evaluation shows current large language models perform poorly, with the best model scoring 43.5% overall and only 17.6% under substitution‑resistant conditions, indicating the benchmark’s difficulty and realism.

By Linyang He, Qiyao Yu, Hanze Dong, Baohao Liao, Xinxing Xu, Micah Goldblum, Jiang Bian, Nima Mesgarani
arXiv AI
Sep 24

Math Reasoning in LLMs is Organized by Approach, Not Topic

The paper argues that large language models (LLMs) organize their internal mathematical reasoning by reusable reasoning approaches rather than by the benchmark topics they are tested on. Using a generation‑replay protocol, the authors extract activation‑importance signatures from eight models across five math sources, cluster these signatures, and find that the resulting groups align more closely with reasoning approaches than with topics. The study shows that changing the requested reasoning approach shifts cluster assignments, while paraphrasing the prompt does not, underscoring the primacy of approach over topic in LLM reasoning.

By Sajad Goudarzi, Samaneh Zamanifard, Moloud Nasiri, Hamed Rahimian