arXiv:2606. 26728v1 Announce Type: new Abstract: Scientific discovery is fundamentally an optimization problem, defined by a vast "state space" of theories and experiments, and an evaluation criterion based on quality, novelty, and validity.
By Yuan-Hang Zhang, Chesson Sipling, Massimiliano Di Ventra
arXiv:2606. 02863v1 Announce Type: new Abstract: AI-Driven Research Systems (ADRS) -- systems coupling LLMs with automated evaluation to discover algorithms, proofs, and designs -- are being optimized and adopted across domains, but the tools to analyze them have not kept pace.
By Marquita Ellis, Paul Castro
arXiv:2606. 10587v1 Announce Type: cross Abstract: Large language models (LLMs) are on the rise for accelerating scientific discovery, most recently in advanced tasks such as generating valid scientific hypotheses.
By Haorui Wang, Parshin Shojaee, Kazem Meidani, Kunyang Sun, Jos\'e Miguel Hern\'andez-Lobato, Teresa Head-Gordon, Jiajun He, Chandan K. Reddy, Chao Zhang, Yuanqi Du
The paper introduces PRISMS, a framework that uses expert pairwise rankings of varying fidelity to curate scientific designs without relying on data-intensive regression models. By escalating queries from lower- to higher-fidelity rankers based on Fisher-information, PRISMS improves discovery recall and reduces the number of screening rounds compared to regression-only and non‑escalated ranking methods. In optimization tasks, PRISMS outperforms Bayesian optimization by achieving higher hypervolume.
By Kevin Tirta Wijaya, Alston Lo, Michael Sun, Wojciech Matusik, Vahid Babaei
arXiv:2604. 19341v2 Announce Type: replace-cross Abstract: Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions.
By Haotian Ye, Haowei Lin, Jingyi Tang, Yizhen Luo, Rahul Thapa, Caiyin Yang, Chang Su, Rui Yang, Ruihua Liu, Rundao Li, Zeyu Li, Pengwei Sun, Chong Gao, Dachao Ding, Guangrong He, Miaolei Zhang, Lina Sun, Wenyang Wang, Yuchen Zhong, Zhuohao Shen, Puheng Li, Pan Lu, Bianxiao Cui, Di He, Jianzhu Ma, Junfeng Li, Hexi Baoyin, Yejin Choi, Stefano Ermon, Xiaowen Chu, Tongyang Li, Yuzhi Xu, James Zou
arXiv:2602. 06448v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based scientific agents have accelerated scientific discovery, yet they often suffer from significant inefficiencies due to adherence to fixed initial priors.
By Yingming Pu, Tao Lin, Hongyu Chen
The paper introduces SCILAWS-BENCH, a benchmark for evaluating large language models (LLMs) on scientific law discovery. It contains 118 problems from 381 scientific papers, covering 291 candidate laws and about 8 million real data points across six disciplines. The benchmark offers two settings: SCILAWS-REAL, where models must propose laws from fixed real observations, and SCILAWS-PARALLEL, where models actively query synthetic worlds to recover hidden laws.
By Yiming Huang, Ziche Liu, Zhuohang Wu, Yiqian Wang, Junxia Cui, Xinkai Zou, Linjun Mao, Nan Huang, Naicheng Yu, Kaijie Zhu, Yue Ma, Kun Zhou, Letian Peng, Jingbo Shang
The paper introduces HorizonMath, a benchmark of 113 largely unsolved mathematical problems across eight domains, paired with an open-source framework for automated verification. It focuses on the generator‑verifier gap, targeting problems that are hard to discover but easy to verify computationally, thereby avoiding costly formal proof verification or manual review. Using this framework, the authors found six novel solutions—three each from GPT‑5.4 Pro and GPT‑5.6 Sol—demonstrating that current models can contribute to mathematical research, while most state‑of‑the‑art models score below 10%.
By Erik Y. Wang, Sumeet R. Motwani, James V. Roggeveen, Eliot Hodges, Dulhan Jayalath, Charles London, Kalyan Ramakrishnan, Jakob Foerster, Cheng Zhang, Flaviu Cipcigan, Philip Torr, Alessandro Abate
arXiv:2607. 28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery.
By Vaibhava Lakshmi Ravideshik, Mayank Kejriwal
arXiv:2512. 03476v3 Announce Type: replace-cross Abstract: Progress in computational science depends on complex numerical workflows that must faithfully encode physical laws, yet translating conceptual insight into reliable code remains a major bottleneck.
By Juan Diego Toscano, Daniel T. Chen, George Em Karniadakis
arXiv:2607. 24647v1 Announce Type: new Abstract: AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks.
By Haiqian Yang, Yuan Cao
D3‑Gym is the first automatically constructed dataset that provides verifiable environments for scientific data‑driven discovery, comprising 565 tasks from 239 real scientific repositories across four disciplines. Each task includes a natural‑language instruction, an executable environment with pre‑installed dependencies, dataset previews, a reference solution, and an automatically synthesized evaluation script that achieves 87.5% agreement with human‑annotated gold standards. Training on trajectories sampled from D3‑Gym consistently improves Qwen3 models on ScienceAgentBench, and the platform also serves as a testbed for studying agentic optimization loops such as Autoresearch on real scientific workflows.
By Hanane Nour Moussa, Yifei Li, Zhuoyang Li, Yankai Yang, Cheng Tang, Tianshu Zhang, Nesreen K. Ahmed, Ali Payani, Ziru Chen, Huan Sun