arXiv AI

Structured Scaling of AI Discovery Across Diverse Scientific Domains

arXiv:2604. 19341v2 Announce Type: replace-cross Abstract: Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions.

arXiv AI
Sep 1

Test-Time Scaling for Scientific Equation Discovery

The paper investigates Test‑Time Scaling (TTS) for large language models (LLMs) in the context of automated scientific equation discovery, an open‑ended task where models iteratively search candidate equations using observed data for feedback. It frames equation discovery as a unified iterative search that encompasses Best‑of‑N, sequential refinement, tree search, and evolutionary methods, and studies how compute allocation—particularly search width—affects performance under fixed budgets. Experiments on the LLM‑SRBench dataset show that increasing search width with more compute improves results, while other factors like population‑branching split and controller choice have smaller impacts, indicating that controlling exploration versus exploitation is key to scaling LLM‑based equation discovery.

By Haowei Lin, Hubert Lim, Xiangyu Wang, Letian Huang, Di He
arXiv AI
Sep 25

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

ExplorationBench is a new benchmark designed to evaluate AI systems’ ability to conduct scientific exploration in verifiable alien worlds. It comprises two sandbox environments—AlienCode and AlienLogic—each containing discovery targets, tasks, flawed manuals, and tool‑call schemas that force systems to formulate hypotheses, design experiments, and iterate on results. Ten AI systems were tested, revealing that while the best performers can learn and apply unfamiliar rules, their progress varies across exploration trajectories and can even regress with continued exploration.

By Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Maxm Pan
Hugging Face Trending Papers
Sep 24

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

ExplorationBench is a benchmark designed to evaluate AI systems’ ability to conduct scientific exploration in verifiable alien worlds, where rules are executable and can be precisely checked. It consists of two sandboxes—AlienCode and AlienLogic—each offering discovery targets, tasks, flawed manuals, environmental feedback, and tool‑call schemas. The benchmark tests whether systems can generate new hypotheses, design experiments, and iterate on results, rather than merely recalling pre‑trained knowledge, and finds that top performers can acquire and apply unfamiliar rules, though performance varies across exploration trajectories.

arXiv Machine Learning
Sep 21

MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery

MOSAIC‑SR is a new symbolic regression method that combines a pretrained Transformer with search‑based refinement. The Transformer generates multiple initial sketches, which seed searches that jointly recover equation structure and constants using scale‑aware optimization and symbolic repair. On the SRSD‑Feynman dataset and six other benchmarks, MOSAIC‑SR achieves the highest symbolic solution rate and ranks among the top two in predictive accuracy, even when irrelevant dummy variables are present.

By Peiyi Zheng, Yanming Kang, Hans De Sterck, Giang Tran
arXiv AI
Aug 14

PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research

arXiv:2512. 19799v2 Announce Type: replace Abstract: Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain expertise, long-horizon reasoning, and reliable numerical computation.

By Tingjia Miao, Wenkai Jin, Jinxin Tan, Muhua Zhang, Xianghe Pang, Zexi Liu, Yuwen Du, Tian Jin, Tu Guo, Zhengliang Zhang, Jingkun Liu, Yuelin Hu, Jiejun Zhang, Yunjie Huang, Yuhan Wang, Wenbo Li, Yinuo Gao, Shuo Chen, Rui Ye, Yuzhi Zhang, Linfeng Zhang, Kun Chen, Wei Wang, Weinan E, Siheng Chen
arXiv Machine Learning
Aug 27

InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance

InsightSR is a new framework that integrates Large Language Models (LLMs) with the PySR genetic programming engine to refine symbolic regression search spaces. It employs two LLM-guided pathways: a Semantic Seed Pathway that generates dimensionally consistent functional skeletons, and a Structural Feature Pathway that suggests nonlinear feature transformations. Over successive iterations, these pathways expand the input space and shift the search toward shallow, semantically informed trees, with a feedback loop that evaluates and refines candidate features. The method achieves a 95% exact recovery rate on the Feynman benchmark and 80.18% accuracy on the LLM-SRBench LSR-Transform task, outperforming existing genetic programming and neural-symbolic approaches while preserving strong out-of-distribution generalization.

By Yating Ling, Wenjing Cun, Zhitang Chen