arXiv AI By Ke Zhang, Yankang Liu, Roya Zandi, Maziar Raissi

ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction

Read the original on arXiv AI →

ToolGate is an executable acceptance pipeline designed to streamline the creation of scientific benchmarks that rely on specialist software. It evaluates each model-generated item through three gates: (1) an executable solution script must reproduce the proposed answer, (2) a randomized no‑tool screen rejects items solvable without the software, and (3) a tool‑using agent must solve the item within a time limit. In a FEniCSx instantiation, 500 generation attempts produced 128 unique, verified benchmark items after successive filtering.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

Can Coding Agents Reproduce Findings in Computational Materials Science?

The paper introduces AutoMat, a benchmark designed to test large language model (LLM) coding agents on their ability to reproduce claims from computational materials science. AutoMat presents three challenges: reconstructing underspecified procedures, navigating specialized toolchains, and assessing whether the evidence supports a claim. Experiments show that current LLM agents achieve low success rates, with the best setting reaching only 53%, and failures stem mainly from incomplete procedures, methodological deviations, and execution fragility.

By Ziyang Huang, Yi Cao, Ali K. Shargh, Jing Luo, Ruidong Mei, Mohd Zaki, Zhan Liu, Wyatt Bunstine, William Jurayj, Somdatta Goswami, Tyrel McQueen, Michael Shields, Jaafar El-Awady, Paulette Clancy, Benjamin Van Durme, Nicholas Andrews, William Walden, Daniel Khashabi