arXiv Machine Learning

When Do Biological Reasoning Models Use Their Biological Inputs?

The study evaluates whether biological reasoning models actually use their biological inputs by testing six models on DNA, protein, and single‑cell tasks. By perturbing one biological input while keeping others fixed, the authors find that many models (e.g., Evo2, ESM3, BioReason, BioReason‑Pro) rely primarily on textual information, with minimal impact from the biological representations. In contrast, models like ChatNT, Prot2Text‑V2, CellWhisperer, and Cell2Sentence‑Scale show greater dependence on their biological inputs, yet overall accuracy gains do not consistently reflect increased biological input contribution.

arXiv Machine Learning
Sep 22

Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

The paper introduces OmicsBench, a new reasoning benchmark for multi‑omics sequences that includes 1,160 expert‑validated questions across DNA regulation, RNA processing, and protein function tasks, requiring traceable evidence chains. Evaluation of 17 large language models shows that scientific LLMs, while more accurate in classification, often lack valid evidence, suggesting shortcut learning. To address this, the authors propose tool‑augmented on‑policy distillation (TA‑OPD), a post‑training method that improves both evidence grounding and predictive performance across five Qwen3.5 models of varying sizes.

By Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong
arXiv Machine Learning
Jun 2

OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction

arXiv:2510. 17532v2 Announce Type: replace-cross Abstract: Predicting cancer treatment outcomes requires models that are both accurate and interpretable, particularly in the presence of heterogeneous clinical data.

By Raghu Vamshi Hemadri, Geetha Krishna Guruju, Kristi Topollai, Anna Ewa Choromanska
arXiv Computation and Language
Sep 14

HypoKG: Evidence-Disciplined Biomedical Hypothesis Generation Beyond Endpoint Knowledge

The paper introduces HypoKG, a unified biochemical knowledge graph built from KEGG, Rhea, and UniProt, and uses it to benchmark 13,200 biomedical hypotheses generated by six large language models (LLMs). By varying the biological information provided—source enzyme only, full biological path, or source and disease endpoint—the study finds that LLMs produce higher-scoring hypotheses when given minimal information, but these are less evidence‑grounded. When supplied with the full biological path, the models generate hypotheses that align more closely with known mechanistic relationships, a phenomenon the authors term evidence‑disciplined reasoning, which is confirmed by shuffling intermediate path steps. "whyItMatters":"The study demonstrates that knowledge graphs can both uncover novel disease–enzyme pairs and guide LLMs to reason more accurately from evidence, improving the reliability of AI‑generated biomedical hypotheses."

By Dominic Okonkwo, Adetayo Okunoye, Ismailcem Budak Arpinar
arXiv AI
4d ago

OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells

OmniVCBench is a figure‑centric, source‑traceable benchmark designed to evaluate the interpretation component of Artificial Intelligence Virtual Cells (AIVCs). It comprises 6,077 curated question–answer pairs drawn from scientific figures and experimental contexts, organized into three scientific reasoning tasks that mirror the AIVC Predict–Explain–Discover agenda. The benchmark also introduces AIVC‑Judge, a task‑conditioned MLLM‑as‑a‑judge framework with reference‑aware rubrics, and a Model‑Derived Hard‑Negative Mining strategy to generate multiple‑choice distractors for efficient evaluation.

By Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
Jun 18

Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment

arXiv:2606. 18703v1 Announce Type: new Abstract: Pretrained biological language models expose per-token probability distributions through masked-token prediction, providing the likelihood interface central to sequence design, variant scoring, and mechanistic interpretation.

By Yanjun Shao, Yundi Chen, Yashvi Patel, Aurelien Pelissier, Mar\'ia Rodr\'iguez Mart\'inez