arXiv:2607. 00436v1 Announce Type: new Abstract: Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex.
By Ke Zhang, Sahchit Chundur, Mohammad Javad Qomi, Maziar Raissi
The paper introduces an experimental model‑class revision framework that jointly proposes structural edits to a hypothesis space and diagnostic experiments to test those edits. By coupling a class‑level distinguishability objective with anytime‑valid sequential evidence, the method only revises the model class after the current one is rejected. On 400 controlled dynamical environments, the approach achieves 89.5% exact recovery with 32 experiments, outperforming baselines and transferring well to unseen mechanisms, library insufficiency detection, and other benchmark tasks.
By SiYuan Ma, Albert Gao, Chunzheng Zhu, Xin Yan, Wenlong Zhang, Wenxin Zhang, Luqi Gong, Tianlin Li, Qixin Zhang
ToolGate is an executable acceptance pipeline designed to streamline the creation of scientific benchmarks that rely on specialist software. It evaluates each model-generated item through three gates: (1) an executable solution script must reproduce the proposed answer, (2) a randomized no‑tool screen rejects items solvable without the software, and (3) a tool‑using agent must solve the item within a time limit. In a FEniCSx instantiation, 500 generation attempts produced 128 unique, verified benchmark items after successive filtering.
By Ke Zhang, Yankang Liu, Roya Zandi, Maziar Raissi
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
By Yu Fu, Yongqi Kang, Yong Zhao
arXiv:2607. 05682v1 Announce Type: new Abstract: LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect.
By Yufeng Wang
arXiv:2606. 13535v1 Announce Type: cross Abstract: Particle physics collider experiments provide Rivet routines as part of the analysis preservation strategy for model-independent measurements.
By Antonio J. Costa, Caterina Doglioni, Christian G\"utschow, Andrew D. Pilkington, Sukanya Sinha