arXiv AI

EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

arXiv:2606. 13602v1 Announce Type: new Abstract: We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis.

arXiv AI
Sep 11

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

OpenDiscoveryTrace is a public dataset of 558 complete AI scientific agent trajectories that records the reasoning process—thoughts, tool calls, observations, errors, revision triggers, and confidence—across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. The dataset includes seven models (three frontier models and four open‑weight models) and 60 live‑retrieval variants, providing a balanced view of performance and error patterns. Pilot analysis shows that process traces reveal behavioral differences invisible to output‑only evaluation, such as differing error rates and types among frontier models.

By Aayam Bansal, Keertan Balaji
arXiv AI
2d ago

Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

The study evaluates whether detailed, profession‑specific system prompts improve performance on scientific tasks. Using an open‑source corpus of 503 agent profiles and Gemini 3.8 Flash, the authors compared matched profiles to four control prompts across nine text‑based science benchmarks and a tool‑using bioinformatics benchmark. Results show no consistent accuracy gains; matched profiles actually increased token usage and cost, and in some cases reduced success rates, with only a minor advantage in one benchmark likely due to prompt length rather than domain expertise.

By Timothy Kassis
arXiv AI
Sep 4

Bioinfoysis Technical Report

The Bioinfoysis Technical Report introduces a multi‑agent harness designed to improve long‑horizon bioinformatics tasks by maintaining persistent, artifact‑grounded analysis runs. It combines global planning with step‑wise, evidence‑driven replanning, ensuring intermediate results are tied to responsible agents and preventing stale evidence reuse. The system was evaluated on BixBench and LAB‑Bench 2, achieving state‑of‑the‑art accuracy and demonstrating that reliable bioinformatics automation relies on robust planning, execution, memory, and evidence flow.

By Qingyang Shao, Xin Zhang, Zhouyang Yuan, Xianying Chen, Yujia Xiang, Zihao Yang, Tong Ye, Yangqi Zhang, Jiakang Xu, Xiaoqing Yan, Xuan Luo, Keyi Li, Enci Fan, Kai Kang, Zhuohan Liu, Xingyu Jin, Chunran Teng, Tao Li, Xinyu Lv, Minghui Wang, Wenfeng Li, Yidan Gao, Siyu Liu, Mingrui Luo, Zhu Liang, Guanren Qiao, Zhiping Xu
arXiv Machine Learning
5d ago

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola
arXiv AI
Jun 16

Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering

arXiv:2606. 16890v1 Announce Type: cross Abstract: Aggregate accuracy benchmarks conceal a systematic structure in how large language models fail at electronic health record (EHR) question answering: questions requiring more inferential steps produce disproportionately more errors.

By Sanjay Basu