arXiv AI

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

arXiv:2601. 12805v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks.

arXiv Machine Learning
Sep 22

Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

The paper introduces OmicsBench, a new reasoning benchmark for multi‑omics sequences that includes 1,160 expert‑validated questions across DNA regulation, RNA processing, and protein function tasks, requiring traceable evidence chains. Evaluation of 17 large language models shows that scientific LLMs, while more accurate in classification, often lack valid evidence, suggesting shortcut learning. To address this, the authors propose tool‑augmented on‑policy distillation (TA‑OPD), a post‑training method that improves both evidence grounding and predictive performance across five Qwen3.5 models of varying sizes.

By Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong
arXiv AI
Jun 4

BRAINCELL-AID: An Agentic AI Created Brain Cell Type Resource for Community Annotation

arXiv:2510. 17064v4 Announce Type: replace Abstract: Single-cell RNA sequencing has transformed our ability to identify diverse cell types and their transcriptomic signatures.

By Rongbin Li, Wenbo Chen, Zhao Li, Rodrigo Munoz-Castaneda, Jinbo Li, Neha S. Maurya, Arnav Solanki, Huan He, Hanwen Xing, Meaghan Ramlakhan, Zachary Wise, Nelson Johansen, Zhuhao Wu, Hua Xu, Michael Hawrylycz, W. Jim Zheng
arXiv AI
Aug 3

ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics

arXiv:2603. 11872v3 Announce Type: replace-cross Abstract: Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain opaque to natural language.

By Omar Coser
arXiv AI
Sep 18

BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research

BioPhys-Bridge is a newly released benchmark dataset designed to evaluate language models on evidence‑grounded scientific reasoning within biophysical literature. Each of its 500 cases includes evidence blocks, stable IDs, quantitative values, units, equations, assumptions, mechanisms, and next‑step decisions, covering six biological domains and nine physical model families. The dataset enforces strict quality gates and has already been evaluated against several models, with DeepSeek‑V4‑Flash achieving the highest evidence‑ID F1 score of 0.360.

By Qingyang Xu
arXiv Computation and Language
Sep 14

HypoKG: Evidence-Disciplined Biomedical Hypothesis Generation Beyond Endpoint Knowledge

The paper introduces HypoKG, a unified biochemical knowledge graph built from KEGG, Rhea, and UniProt, and uses it to benchmark 13,200 biomedical hypotheses generated by six large language models (LLMs). By varying the biological information provided—source enzyme only, full biological path, or source and disease endpoint—the study finds that LLMs produce higher-scoring hypotheses when given minimal information, but these are less evidence‑grounded. When supplied with the full biological path, the models generate hypotheses that align more closely with known mechanistic relationships, a phenomenon the authors term evidence‑disciplined reasoning, which is confirmed by shuffling intermediate path steps. "whyItMatters":"The study demonstrates that knowledge graphs can both uncover novel disease–enzyme pairs and guide LLMs to reason more accurately from evidence, improving the reliability of AI‑generated biomedical hypotheses."

By Dominic Okonkwo, Adetayo Okunoye, Ismailcem Budak Arpinar
arXiv Machine Learning
1d ago

When Do Biological Reasoning Models Use Their Biological Inputs?

The study evaluates whether biological reasoning models actually use their biological inputs by testing six models on DNA, protein, and single‑cell tasks. By perturbing one biological input while keeping others fixed, the authors find that many models (e.g., Evo2, ESM3, BioReason, BioReason‑Pro) rely primarily on textual information, with minimal impact from the biological representations. In contrast, models like ChatNT, Prot2Text‑V2, CellWhisperer, and Cell2Sentence‑Scale show greater dependence on their biological inputs, yet overall accuracy gains do not consistently reflect increased biological input contribution.

By Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
arXiv AI
4d ago

OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells

OmniVCBench is a figure‑centric, source‑traceable benchmark designed to evaluate the interpretation component of Artificial Intelligence Virtual Cells (AIVCs). It comprises 6,077 curated question–answer pairs drawn from scientific figures and experimental contexts, organized into three scientific reasoning tasks that mirror the AIVC Predict–Explain–Discover agenda. The benchmark also introduces AIVC‑Judge, a task‑conditioned MLLM‑as‑a‑judge framework with reference‑aware rubrics, and a Model‑Derived Hard‑Negative Mining strategy to generate multiple‑choice distractors for efficient evaluation.

By Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li, Guoren Wang
arXiv AI
Aug 18

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

arXiv:2608. 16419v1 Announce Type: cross Abstract: Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces.

By Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu
arXiv AI
Jul 13

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.

By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung