SciR: A Controllable Benchmark for Scientific Reasoning in LLMs
arXiv:2606. 13020v1 Announce Type: new Abstract: Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction.
arXiv:2606. 13020v1 Announce Type: new Abstract: Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction.
arXiv:2603. 05290v2 Announce Type: replace Abstract: Large language models (LLMs) achieve promising performance, yet their ability to reason remains poorly understood.
arXiv:2607. 07391v1 Announce Type: new Abstract: Mathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasoning with tools, retrieval, and long-horizon dialogue.
arXiv:2608. 03600v1 Announce Type: new Abstract: Partial differential equations (PDEs) become actionable in science and engineering not as isolated formulae, but as executable workflows that connect modelling assumptions, governing equations, numerical solvers, diagnostics, and decisions.
arXiv:2510. 12171v2 Announce Type: replace Abstract: Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied.
arXiv:2607. 20926v1 Announce Type: new Abstract: Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources.
arXiv:2606.15872v2 Announce Type: replace Abstract: Frontier scientific reasoning remains a major challenge for large language models (LLMs), where even the strongest commercial systems fall short of...
SciMIF is a new benchmark that evaluates how well multimodal large language models (MLLMs) can follow complex scientific instructions. It is built on an analysis of 22 tasks across five scientific fields and introduces a taxonomy of 10 constraint groups that capture both general and discipline‑specific requirements. Experiments show large performance gaps between fields—chemistry is hardest—and that larger models do not necessarily improve constraint adherence, especially for fine‑grained, knowledge‑heavy instructions.
arXiv:2509. 21028v4 Announce Type: replace Abstract: We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language models (LLMs).
arXiv:2607. 24459v1 Announce Type: new Abstract: Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems.
arXiv:2603.18614v2 Announce Type: replace Abstract: Tool-augmented large language models (LLMs) must tightly couple multi-step reasoning with external actions, yet existing benchmarks often confound...
arXiv:2509. 03059v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforcement Learning with Verifiable Reward (RLVR), particularly in domains like mathematics and programming, where ground-truth correctness can be automatically evaluated.