arXiv AI By Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke

Rules to Tools: Executable Checks for LLM Agents in Scientific Computing

Read the original on arXiv AI →

The paper introduces Rules to Tools (R2T), a system that provides executable checks for scientific coding agents to verify compliance with public scientific requirements. In experiments across multiple task cohorts, agents using R2T’s prepared checks achieved high repair success rates—26/30 with text and 29/30 with checks—while also demonstrating varying task preferences and cost trade‑offs. The study quantifies how tool‑enabled checks influence repair outcomes and agent‑side resource usage in scientific computing contexts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction

ToolGate is an executable acceptance pipeline designed to streamline the creation of scientific benchmarks that rely on specialist software. It evaluates each model-generated item through three gates: (1) an executable solution script must reproduce the proposed answer, (2) a randomized no‑tool screen rejects items solvable without the software, and (3) a tool‑using agent must solve the item within a time limit. In a FEniCSx instantiation, 500 generation attempts produced 128 unique, verified benchmark items after successive filtering.

By Ke Zhang, Yankang Liu, Roya Zandi, Maziar Raissi
arXiv AI
Sep 11

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

OpenDiscoveryTrace is a public dataset of 558 complete AI scientific agent trajectories that records the reasoning process—thoughts, tool calls, observations, errors, revision triggers, and confidence—across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. The dataset includes seven models (three frontier models and four open‑weight models) and 60 live‑retrieval variants, providing a balanced view of performance and error patterns. Pilot analysis shows that process traces reveal behavioral differences invisible to output‑only evaluation, such as differing error rates and types among frontier models.

By Aayam Bansal, Keertan Balaji