arXiv:2604. 17621v2 Announce Type: replace Abstract: Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based reasoning over that universe, a phenomenon we term "the tip of the iceberg.
By Xiao Zhang, Qianru Meng, Yongjian Chen, Yumeng Wang, Johan Bos
arXiv:2605. 19723v2 Announce Type: replace-cross Abstract: Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems.
By Husnain Amjad, Raja Khurram Shahzad, Aamir Shahzad, Mehwish Fatima
arXiv:2610. 02191v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions.
By Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez, Mohit Bansal, Zhengzhong Tu
arXiv:2606. 10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processes of real human students remains under-examined.
By Yiteng Mao, Kenan Xu, Yijia Lyu, Wenhao Li, Jianlong Chen, Xiangfeng Wang
arXiv:2608. 07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures.
By Shibo Chu, Yuze Liu, Tiehua Zhang, Zhishu Shen, Lianghua He, Haofen Wang, Zhijun Ding
The paper introduces the Structure-Internalized Rule Language Model (SIRLM) to improve Knowledge Graph Reasoning (KGR) by addressing the mismatch between KG structural context and Large Language Model (LLM) parametric knowledge. SIRLM centers on a Structure-Internalized Rule Generator (SIRG) that uses in-context learning, a structural relation memory, a KG tokenizer, and a neuro-symbolic reasoner to generate structural rules and provide faithful rule-execution feedback. Experiments on 36 datasets against 17 state‑of‑the‑art KGR methods show that SIRLM achieves significant performance gains.
By Xingrui Zhuo, Jiapu Wang, Manzong Huang, Gongqing Wu, Xindong Wu
arXiv:2607. 01431v1 Announce Type: cross Abstract: We introduce ISOSCI, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge retrieval in LLM evaluation.
By Samir Abdaljalil, Erchin Serpedin, Hasan Kurban
arXiv:2606. 27378v1 Announce Type: cross Abstract: We introduce an axiomatic evaluation framework for latent thought representations in LLMs, comprising metrics that are independent of downstream benchmark scores and reveal representational failures that benchmark accuracy masks.
By Fahd Seddik, Fatemeh Fard
arXiv:2604. 27540v2 Announce Type: replace Abstract: Scientific reasoning rarely stops at what is directly observable; it often requires uncovering hidden structure from data.
By Chaemin Jang, Woojin Park, Hyeok Yun, Dongman Lee, Jihee Kim
LiveMathematicianBench is a dynamic multiple‑choice benchmark for research‑level mathematical reasoning, built from recent arXiv papers published after model training cutoffs. It introduces a thirteen‑category logical taxonomy of theorem types and uses a proof‑sketch‑guided distractor pipeline to create plausible but invalid answer choices, enhancing sensitivity to genuine reasoning. Evaluation shows current large language models perform poorly, with the best model scoring 43.5% overall and only 17.6% under substitution‑resistant conditions, indicating the benchmark’s difficulty and realism.
By Linyang He, Qiyao Yu, Hanze Dong, Baohao Liao, Xinxing Xu, Micah Goldblum, Jiang Bian, Nima Mesgarani
arXiv:2607. 08393v1 Announce Type: new Abstract: Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks.
By Lu Dai, Ziyang Rao, Yili Wang, Hanqing Wang, Hao Liu, Hui Xiong
The paper argues that large language models (LLMs) organize their internal mathematical reasoning by reusable reasoning approaches rather than by the benchmark topics they are tested on. Using a generation‑replay protocol, the authors extract activation‑importance signatures from eight models across five math sources, cluster these signatures, and find that the resulting groups align more closely with reasoning approaches than with topics. The study shows that changing the requested reasoning approach shifts cluster assignments, while paraphrasing the prompt does not, underscoring the primacy of approach over topic in LLM reasoning.
By Sajad Goudarzi, Samaneh Zamanifard, Moloud Nasiri, Hamed Rahimian