arXiv AI By Kautik Mandve, Dileepa Fernando

Explore, Then Commit: Measurement-Efficient Scientific Law Discovery with Language Models

Read the original on arXiv AI →

The paper presents an explore‑then‑commit protocol that uses a large language model to generate hypotheses, a programmatic planner to collect measurements, and a fresh prompt to synthesize a scientific law from fixed observations. In 576 NewtonBench trials across 12 physics modules, the protocol—especially when interpreter‑enabled planners are used—reduces the number of measurements needed and improves root‑mean‑squared logarithmic error for both GPT‑4.1‑mini and GPT‑4.1. The study demonstrates measurement savings in every module, though it notes that the causal components and generalization beyond noiseless direct‑equation tasks remain unresolved.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
1d ago

Learning to Outgrow a Theory: Experimental Discovery Beyond the Initial Hypothesis Space

The paper introduces an experimental model‑class revision framework that jointly proposes structural edits to a hypothesis space and diagnostic experiments to test those edits. By coupling a class‑level distinguishability objective with anytime‑valid sequential evidence, the method only revises the model class after the current one is rejected. On 400 controlled dynamical environments, the approach achieves 89.5% exact recovery with 32 experiments, outperforming baselines and transferring well to unseen mechanisms, library insufficiency detection, and other benchmark tasks.

By SiYuan Ma, Albert Gao, Chunzheng Zhu, Xin Yan, Wenlong Zhang, Wenxin Zhang, Luqi Gong, Tianlin Li, Qixin Zhang
arXiv AI
Sep 3

ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction

ToolGate is an executable acceptance pipeline designed to streamline the creation of scientific benchmarks that rely on specialist software. It evaluates each model-generated item through three gates: (1) an executable solution script must reproduce the proposed answer, (2) a randomized no‑tool screen rejects items solvable without the software, and (3) a tool‑using agent must solve the item within a time limit. In a FEniCSx instantiation, 500 generation attempts produced 128 unique, verified benchmark items after successive filtering.

By Ke Zhang, Yankang Liu, Roya Zandi, Maziar Raissi
arXiv AI
Jul 8

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

arXiv:2607. 05682v1 Announce Type: new Abstract: LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect.

By Yufeng Wang