arXiv AI By Timothy Kassis

Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

Read the original on arXiv AI →

The study evaluates whether detailed, profession‑specific system prompts improve performance on scientific tasks. Using an open‑source corpus of 503 agent profiles and Gemini 3.8 Flash, the authors compared matched profiles to four control prompts across nine text‑based science benchmarks and a tool‑using bioinformatics benchmark. Results show no consistent accuracy gains; matched profiles actually increased token usage and cost, and in some cases reduced success rates, with only a minor advantage in one benchmark likely due to prompt length rather than domain expertise.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 11

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

OpenDiscoveryTrace is a public dataset of 558 complete AI scientific agent trajectories that records the reasoning process—thoughts, tool calls, observations, errors, revision triggers, and confidence—across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. The dataset includes seven models (three frontier models and four open‑weight models) and 60 live‑retrieval variants, providing a balanced view of performance and error patterns. Pilot analysis shows that process traces reveal behavioral differences invisible to output‑only evaluation, such as differing error rates and types among frontier models.

By Aayam Bansal, Keertan Balaji
arXiv AI
2d ago

Rules to Tools: Executable Checks for LLM Agents in Scientific Computing

The paper introduces Rules to Tools (R2T), a system that provides executable checks for scientific coding agents to verify compliance with public scientific requirements. In experiments across multiple task cohorts, agents using R2T’s prepared checks achieved high repair success rates—26/30 with text and 29/30 with checks—while also demonstrating varying task preferences and cost trade‑offs. The study quantifies how tool‑enabled checks influence repair outcomes and agent‑side resource usage in scientific computing contexts.

By Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam