arXiv AI

Evidence-Informed LLM Beliefs for Continual Scientific Discovery

arXiv:2606. 29182v1 Announce Type: new Abstract: Open-ended scientific discovery with large language models (LLMs) increasingly operates as a long-horizon loop of hypothesis search and verification, where a reward signal guides which hypotheses to test next.

arXiv AI
Sep 7

Evidence Integration in Large Language Models

The paper proposes a distributional theory explaining how large language models (LLMs) incorporate external evidence into their decision-making process. It identifies three key predictions: (1) evidence is more persuasive when it aligns with the model’s prior beliefs, (2) models more readily accept errors from their own internal processes than from external sources, and (3) the same evidence can improve weaker models while harming stronger ones. Extensive experiments across ten million trials, twelve LLMs from four families, and eight domains—including quantum mechanics, physics, genetics, and molecular biology—confirm these predictions and reveal that evidence integration occurs late in the network as a structured sequence of steps rather than through a simple trust metric.

By Sebastien Kawada, Manolis Kellis
arXiv Machine Learning
Aug 18

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

arXiv:2608. 15669v1 Announce Type: new Abstract: Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs.

By Zhongwei Yu, Yan Song, Xue Yan, Anjie Liu, Xingyu Lu, Yihang Chen, Huichi Zhou, Siyuan Guo, Luoyang Sun, Sihan Chen, Xiangning Yu, Jun Wang
arXiv AI
Aug 28

Learning the ARTS of Search for Automated Discovery

The paper introduces Agentic Reasoning for Tree Search (ARTS), a method that uses a reasoning language model to navigate the hypothesis‑experiment space in scientific discovery. Unlike traditional approaches that conflate hypothesis quality with execution quality and prune search logs, ARTS evaluates prior execution logs to distinguish implementation failures from poor hypotheses and selects the next hypothesis to pursue. By employing test‑time training to embed search‑tree knowledge into model weights, ARTS achieves a 15.3% relative improvement over leading algorithms on 22 benchmark tasks and enables smaller models like Qwen3‑4B to match or exceed the performance of larger closed‑source models at lower inference cost.

By Gurusha Juneja, Arnav Kumar Jain, Deepak Nathani, William Yang Wang, Xin Eric Wang
arXiv AI
Aug 3

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

arXiv:2607. 28684v1 Announce Type: new Abstract: Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model is discovering laws from data or merely recalling answers from its training corpus.

By Zhan'ao Yao, Liang Yin, Zhihao Gao, Boxuan Zhang, Xiaoyu Wu, Linjing Li, Rongyan Wang, Tingwei Chen, Youwei Wang, Xiaolin Zhao, Jiahui Shi, Jianjun Liu
arXiv AI
Jul 7

Amortising Bayesian Experimental Design for Sequential Information Gathering in LLMs

arXiv:2607. 03426v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectively across the multi-turn interactions required in sequential decision-making settings.

By Jakob Hartmann, James Harvey, Jhonathan Navott, Erik Y. Wang, Luckeciano C. Melo, Flaviu Cipcigan, Cheng Zhang, Alessandro Abate