arXiv AI By Samir Abdaljalil, Erchin Serpedin, Hasan Kurban

IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs

Read the original on arXiv AI →

arXiv:2607. 01431v1 Announce Type: cross Abstract: We introduce ISOSCI, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge retrieval in LLM evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches

LiveMathematicianBench is a dynamic multiple‑choice benchmark for research‑level mathematical reasoning, built from recent arXiv papers published after model training cutoffs. It introduces a thirteen‑category logical taxonomy of theorem types and uses a proof‑sketch‑guided distractor pipeline to create plausible but invalid answer choices, enhancing sensitivity to genuine reasoning. Evaluation shows current large language models perform poorly, with the best model scoring 43.5% overall and only 17.6% under substitution‑resistant conditions, indicating the benchmark’s difficulty and realism.

By Linyang He, Qiyao Yu, Hanze Dong, Baohao Liao, Xinxing Xu, Micah Goldblum, Jiang Bian, Nima Mesgarani
arXiv AI
3d ago

LSR-Ben: A Logical and Scientific Reasoning Benchmark for Evaluating Process Reward Models

The paper introduces LSR‑Ben, a benchmark designed to evaluate process reward models (PRMs) on scientific and logical reasoning tasks, addressing a gap left by existing math‑focused benchmarks. Experiments on 22 models reveal that PRMs and LLMs perform poorly in non‑mathematical domains, with LLMs tending to over‑identify errors while PRMs tend to overlook them. LSR‑Ben aims to spur research that broadens PRM applicability and improves LLM reasoning.

By Zhouhao Sun, Xuan Zhang, Xiao Ding, Bibo Cai, Li Du, Kai Xiong, Xinran Dai, Fei Zhang, weidi tang, Zhiyuan Kan, Yang Zhao, Bing Qin, Ting Liu
arXiv AI
Jul 15

Rethinking Reward Models for Multi-Domain Test-Time Scaling

arXiv:2510. 00492v3 Announce Type: replace Abstract: The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning from flawed logic.

By Dong Bok Lee, Seanie Lee, Sangwoo Park, Minki Kang, Jinheon Baek, Dongki Kim, Dominik Wagner, Jiongdao Jin, Heejun Lee, Tobias Bocklet, Jinyu Wang, Jingjing Fu, Sung Ju Hwang, Jiang Bian, Lei Song
arXiv AI
2d ago

Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI

The paper introduces a pipeline that automatically creates ontology‑grounded multiple‑choice question benchmarks for evaluating large language models (LLMs) on logical reasoning tasks in scientific AI. By using OWL 2 ontologies, correct answers are guaranteed by design and distractors are generated and formally verified as incorrect through an OWL reasoner. Experiments on three ontologies—Pizza, PMDco, and DOID—yielded 112, 2,491, and 15,216 MCQs, respectively, with high natural‑language quality and challenging zero‑shot performance for six LLMs.

By Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler