arXiv AI

FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games

arXiv:2606. 04751v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks.

arXiv AI
Jul 14

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.

By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
arXiv AI
Jun 11

Can AI Agents Synthesize Scientific Conclusions?

arXiv:2606. 11337v1 Announce Type: new Abstract: Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions.

By Hayoung Jung, Pedro Viana Diniz, Jos\'e Reinaldo Corr\^ea Roveda, Abner Fernandes da Silva, Haeun Jung, Enoch Tsai, Aleksandra Korolova, Manoel Horta Ribeiro
arXiv AI
Sep 7

Evidence Integration in Large Language Models

The paper proposes a distributional theory explaining how large language models (LLMs) incorporate external evidence into their decision-making process. It identifies three key predictions: (1) evidence is more persuasive when it aligns with the model’s prior beliefs, (2) models more readily accept errors from their own internal processes than from external sources, and (3) the same evidence can improve weaker models while harming stronger ones. Extensive experiments across ten million trials, twelve LLMs from four families, and eight domains—including quantum mechanics, physics, genetics, and molecular biology—confirm these predictions and reveal that evidence integration occurs late in the network as a structured sequence of steps rather than through a simple trust metric.

By Sebastien Kawada, Manolis Kellis
arXiv AI
Sep 25

LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches

LiveMathematicianBench is a dynamic multiple‑choice benchmark for research‑level mathematical reasoning, built from recent arXiv papers published after model training cutoffs. It introduces a thirteen‑category logical taxonomy of theorem types and uses a proof‑sketch‑guided distractor pipeline to create plausible but invalid answer choices, enhancing sensitivity to genuine reasoning. Evaluation shows current large language models perform poorly, with the best model scoring 43.5% overall and only 17.6% under substitution‑resistant conditions, indicating the benchmark’s difficulty and realism.

By Linyang He, Qiyao Yu, Hanze Dong, Baohao Liao, Xinxing Xu, Micah Goldblum, Jiang Bian, Nima Mesgarani