Structured Scaling of AI Discovery Across Diverse Scientific Domains
arXiv:2604. 19341v2 Announce Type: replace-cross Abstract: Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions.
The paper investigates Test‑Time Scaling (TTS) for large language models (LLMs) in the context of automated scientific equation discovery, an open‑ended task where models iteratively search candidate equations using observed data for feedback. It frames equation discovery as a unified iterative search that encompasses Best‑of‑N, sequential refinement, tree search, and evolutionary methods, and studies how compute allocation—particularly search width—affects performance under fixed budgets. Experiments on the LLM‑SRBench dataset show that increasing search width with more compute improves results, while other factors like population‑branching split and controller choice have smaller impacts, indicating that controlling exploration versus exploitation is key to scaling LLM‑based equation discovery.
arXiv:2604. 19341v2 Announce Type: replace-cross Abstract: Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions.
arXiv:2511. 22651v2 Announce Type: replace-cross Abstract: Optimization methods have long advanced many fields, yet they struggle when faced with design problems where the search space and design parameters are difficult to define.
arXiv:2606. 29082v1 Announce Type: cross Abstract: Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture?
arXiv:2608.30395v1 Announce Type: new Abstract: As pretraining scaling laws approach saturation, Test-Time Scaling (TTS) has emerged as an important direction for improving reasoning by allocating in...
arXiv:2602. 09574v2 Announce Type: replace-cross Abstract: Tree-search decoding is an effective form of test-time scaling for large language models (LLMs), but real-world deployment often imposes a fixed per-query token budget that varies across settings.
arXiv:2607. 04156v1 Announce Type: new Abstract: Scientific equation discovery must combine broad domain priors with strict numerical testing.
arXiv:2605. 25246v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines.
arXiv:2606. 02863v1 Announce Type: new Abstract: AI-Driven Research Systems (ADRS) -- systems coupling LLMs with automated evaluation to discover algorithms, proofs, and designs -- are being optimized and adopted across domains, but the tools to analyze them have not kept pace.
arXiv:2606. 18284v1 Announce Type: cross Abstract: The limiting resource for training agents via reinforcement learning (RL) is increasingly frontier task supply: valid, solvable tasks just difficult enough to train the current model.
arXiv:2602. 10576v2 Announce Type: replace-cross Abstract: Symbolic regression aims to distill mathematical equations from observational data.
arXiv:2609.25575v1 Announce Type: cross Abstract: Fine-tuned Large Language Models (LLMs) significantly advance Automated Theorem Proving (ATP), but are often deployed as guiding policies within tree...
arXiv:2606. 01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed sample budget, a fixed refinement loop, a fixed scoring rule, or a fixed search policy decides how compute is spent, leaving the model in charge of solving but not of orchestration.