arXiv AI By Harihara Muralidharan, Reema Baskar, Soo Hee Lee, Tim Proctor, Kenny Workman

EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

Read the original on arXiv AI →

arXiv:2606. 13602v1 Announce Type: new Abstract: We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 11

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

OpenDiscoveryTrace is a public dataset of 558 complete AI scientific agent trajectories that records the reasoning process—thoughts, tool calls, observations, errors, revision triggers, and confidence—across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. The dataset includes seven models (three frontier models and four open‑weight models) and 60 live‑retrieval variants, providing a balanced view of performance and error patterns. Pilot analysis shows that process traces reveal behavioral differences invisible to output‑only evaluation, such as differing error rates and types among frontier models.

By Aayam Bansal, Keertan Balaji
arXiv AI
2d ago

Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

The study evaluates whether detailed, profession‑specific system prompts improve performance on scientific tasks. Using an open‑source corpus of 503 agent profiles and Gemini 3.8 Flash, the authors compared matched profiles to four control prompts across nine text‑based science benchmarks and a tool‑using bioinformatics benchmark. Results show no consistent accuracy gains; matched profiles actually increased token usage and cost, and in some cases reduced success rates, with only a minor advantage in one benchmark likely due to prompt length rather than domain expertise.

By Timothy Kassis