arXiv:2607. 23975v1 Announce Type: new Abstract: Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity.
By Stefan G. Creadore
arXiv:2606. 11208v1 Announce Type: cross Abstract: Biomedical findings often seem to conflict across studies, but many of these differences are context-dependent rather than true contradictions.
By Elias Hossain, Sanjeda Sara Jennifer, Sabera Akter Bushra, Niloofar Yousefi
OpenDiscoveryTrace is a public dataset of 558 complete AI scientific agent trajectories that records the reasoning process—thoughts, tool calls, observations, errors, revision triggers, and confidence—across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. The dataset includes seven models (three frontier models and four open‑weight models) and 60 live‑retrieval variants, providing a balanced view of performance and error patterns. Pilot analysis shows that process traces reveal behavioral differences invisible to output‑only evaluation, such as differing error rates and types among frontier models.
By Aayam Bansal, Keertan Balaji
arXiv:2606. 19245v1 Announce Type: new Abstract: Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions.
By Hannah Le, Ramesh Ramasamy, Alex Urrutia, Mahsa Yazdani, Tim Proctor, Kenny Workman
arXiv:2607. 09349v1 Announce Type: cross Abstract: Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents.
By Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim
The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.
By Yiqi Yao, Miquel Duran-Frigola
arXiv:2609.21859v1 Announce Type: new
Abstract: Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely o...
By Jiacheng Lin, Zifeng Wang, Zheng Chen, Erick Scott, Ziwei Yang, Fanyang Yu, Sheng Zhong, Jimeng Sun
arXiv:2608.25466v2 Announce Type: replace-cross
Abstract: The functional annotation of genes in non-model organisms remains a significant challenge in computational biology, with 20-70% of sequenced...
By Azrin Sultana
TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.
By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
arXiv:2607. 24878v1 Announce Type: cross Abstract: Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranked above it or what evidence a tool-using model examines before making its decision.
By Guiling Guo, Jia Yang, Jiahao Xu, Shuyuan Zheng, Zhonghai Sun, Qiyuan Li
arXiv:2606. 16149v1 Announce Type: new Abstract: Most medical AI systems improve by scaling additional machinery: more fine-tuning data, more agents, and/or larger retrieval databases.
By Minh-Ha Nguyen, Erica Gray, Chih-Ting Yang, Rizwan Hamid, Lingyao Li, Siyuan Ma, Thomas A. Cassini, Cathy Shyr
Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace current safeguards. To assess the biological risks of frontier models, we develop Intern-BioBreaker, a specialized bio-red-teaming model, together with an integrated computational-to-physical framework that couples model-level stress testing with wet-lab validation.